Token-Disentangled Latent Test-Time Scaling for Vision-Language Reasoning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation in multimodal reasoning where global hidden state updates overlook differences in token roles. To this end, this work proposes a token-decoupled test-time scaling framework. Leveraging an entropy-aware routing mechanism, the method disentangles vision-sensitive tokens from high-entropy reasoning tokens and applies targeted feedback to each group. By integrating latent space optimization with anchor regularization, it achieves fine-grained, token-level coordination between perception and reasoning. Experimental results demonstrate that the proposed framework improves macro-average accuracy on models such as Qwen2.5-VL by 2.57% and 1.51%, respectively, compared to standard Chain-of-Thought (CoT) baselines, significantly outperforming existing methods.
📝 Abstract
Latent test-time scaling improves reasoning by refining hidden states during inference, but existing methods typically apply a single scalar reward to all editable latent tokens. For multimodal large language models, this global update ignores that generated tokens play different roles: some are sensitive to visual evidence, while others correspond to uncertain reasoning decisions. We present Token-Disentangled Latent Test-Time Scaling, an inference-time framework that makes latent refinement token-role-aware. Starting from an initial generated trajectory, we optimize a short hidden-state prefix while routing perception-side visual feedback to image-sensitive tokens and reasoning feedback to high-entropy tokens. Tokens selected by neither route are constrained by an anchor regularizer. Across both perception and reasoning benchmarks on Qwen2.5-VL-7B and InternVL3.5-8B, our method lifts macro accuracy over CoT by +2.57 and +1.51 respectively, and outperforms strong output-space test-time scaling baselines under matched decoded-candidate budgets. Code is available at https://github.com/Qwen-Applications/TD-LTTS.
Problem

Research questions and friction points this paper is trying to address.

test-time scaling
vision-language reasoning
multimodal large language models
token disentanglement
latent refinement
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent test-time scaling
Token disentanglement
Vision-language reasoning
Hidden-state optimization
Multimodal large language models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hao-Xuan Ma
Qwen Business Unit of Alibaba, School of Artificial Intelligence, Nanjing University, National Key Laboratory for Novel Software Technology, Nanjing University
Y
Yihao Liu
Qwen Business Unit of Alibaba
Yutao Sun
Yutao Sun
Tsinghua University
Natural Language ProcessingMachine Learning
Y
Yanting Miao
Qwen Business Unit of Alibaba, University of Waterloo
Mengyu Zhou
Mengyu Zhou
Microsoft Research
Data analyticsNatural Language ProcessingNetwork ScienceHuman BehaviorsMobile & Ubiquitous Computing
Y
YiCheng Xiao
Chinese Academy of Sciences
L
Long Chen
The Hong Kong University of Science and Technology
Zhenguo Li
Zhenguo Li
Huawei Noah's Ark Lab, Columbia, CUHK, PKU
machine learninggenerative AIAI for mathematics
Han-Jia Ye
Han-Jia Ye
Nanjing University
Machine LearningData MiningMetric LearningMeta-Learning
X
Xiaoxi Jiang
Qwen Business Unit of Alibaba
G
Guanjun Jiang
Qwen Business Unit of Alibaba