MoVISA: Multi-Token Reasoning for Video Object Segmentation

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insufficient spatiotemporal localization granularity of existing single-token strategies in video multi-object segmentation by proposing the MoVISA framework. Leveraging multimodal large language models (MLLMs), this work pioneers a multi-token reasoning mechanism to replace the conventional single-token paradigm, employing multiple segmentation tokens to represent cross-frame objects and thereby achieving fine-grained alignment between language prompts and spatiotemporal masks. MoVISA significantly enhances both the performance and interpretability of video object segmentation, yielding improvements of 13.2% and 8.4% in the J&F metric on the MeViS and ReVOS benchmarks, respectively.
📝 Abstract
Recent advances in video object segmentation with Multimodal Large Language Model (MLLM) reasoning have demonstrated the effectiveness of using a single textual token, such as SEG, to predict segmentation masks across images and videos. However, we observe that this single-token strategy lacks the granularity required to precisely localize multiple objects across time in video segmentation tasks. To address this limitation, we develop Multi-Token Reasoning for Video Object Segmentation, or MoVISA. MoVISA uses multiple segmentation tokens, such as SEG0 and SEG1, to represent an object across different frames. This design enables more fine-grained alignment between language prompts and spatio-temporal mask predictions, improving both performance and interpretability. On the challenging MeViS, DAVIS17, ReVOS, and Ref-Youtube-VOS benchmarks, our model achieves a 13.2 percent J and F improvement on MeViS and an 8.4 percent J and F improvement on ReVOS. Code and models will be released.
Problem

Research questions and friction points this paper is trying to address.

Video Object Segmentation
Multimodal Large Language Model
Single-Token Strategy
Multi-Object Localization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Token Reasoning
Video Object Segmentation
Multimodal Large Language Model
Spatio-temporal Alignment
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Ruining Zhao
University of Illinois Urbana-Champaign
Ho Kei Cheng
Ho Kei Cheng
University of Illinois Urbana-Champaign
Computer VisionMachine Learning
A
Alexander G. Schwing
University of Illinois Urbana-Champaign