Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding

📅 2026-07-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing approaches struggle to jointly achieve precise visual grounding and complex semantic reasoning, primarily due to the limited expressiveness of textual coordinate prediction and the absence of spatial reference in implicit tokens. This work proposes Mixture-of-Thought-Tokens (Motto), a novel framework that introduces spatially explicit thought tokens to enable visually interpretable perception and constructs context-adaptive dynamic reasoning chains. For the first time, Motto unifies and jointly optimizes dense grounding and high-order reasoning within a single architecture. The method integrates spatially anchored tokenization with fine-tuning of multimodal large language models and demonstrates state-of-the-art performance across diverse free-form referring tasks on the newly introduced PR-Bench benchmark, significantly narrowing the capability gap between perception and reasoning.
📝 Abstract
Multimodal Large Language Models have made great progress in grounding tasks, yet existing methods still struggle to unify precise localization and complex reasoning. For one thing, text-based methods rely on coordinates or index prediction, severely limiting the perceptual capabilities of the model for dense visual objects. Meanwhile, latent token-based methods employ special tokens without inherent spatial references and use a decoding mechanism that lacks thinking steps, weakening high-level reasoning capabilities. Consequently, developing a unified framework that excels in both perception and reasoning remains challenging. To address this, we propose Mixture-of-Thought-Tokens (Motto), a new free-form multimodal grounding method that bridges the perception-reasoning gap, enabling MLLMs to empower diverse, arbitrary grounding queries. Specifically, we introduce Spatially-Grounded Thought Tokenization to explicitly align special tokens with spatial locations for clear spatial correspondence and visual interpretability. We further design a Context-Adaptive Chain-of-Tokens that dynamically switch grounding modes within an interleaved reasoning chain, achieving robust grounding across tasks of varying complexity. In addition, we construct PR-Bench, a new referring expression comprehension benchmark to evaluate the perception-reasoning gap. Extensive experiments demonstrate that Motto achieves state-of-the-art performance across diverse free-form grounding tasks.
Problem

Research questions and friction points this paper is trying to address.

multimodal grounding
perception-reasoning gap
spatial localization
complex reasoning
referring expression comprehension
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Thought-Tokens
Spatially-Grounded Thought Tokenization
Context-Adaptive Chain-of-Tokens
multimodal grounding
perception-reasoning unification
Tianyi Gao
Tianyi Gao
Washington University in St. Louis
Han Fang
Han Fang
TeleAI, China Telecom (中国电信人工智能研究院 TeleAI)
Text-to-imageMLLMVideo-text retrievalFace recognition
T
Tianyi Ding
Xingchen AGI Lab, China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd; Beijing University of Posts and Telecommunications
H
Hao Li
Xingchen AGI Lab, China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd; Shanghai Jiao Tong University
Xin Wei
Xin Wei
Schmidt AI in Science Postdoc, University of Michigan
Natural HazardsAI for GeohazardsResilienceRiskReliability
Hongbo Sun
Hongbo Sun
Institute of Artificial Intelligence, China Telecom(中国电信人工智能研究院,TeleAI), Peking University
Fine-grained visual analysisMulti-modal understandingMachine learning
X
Xiaodong Dong
Xingchen AGI Lab, China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd
Ye Yuan
Ye Yuan
Chinatelecom
computer vison;machine learning
J
Jinglin Xu
University of Science and Technology Beijing
Kongming Liang
Kongming Liang
Beijing University of Posts and Telecommunications
Computer VisionPattern RecognitionMachine Learning
H
Hao Sun
Xingchen AGI Lab, China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd
Jingmin Xin
Jingmin Xin
Xi'an Jiaotong University
Statistical and Array Sensor ArrayPattern Recognition