🤖 AI Summary
This study addresses the limitation in multimodal reasoning where global hidden state updates overlook differences in token roles. To this end, this work proposes a token-decoupled test-time scaling framework. Leveraging an entropy-aware routing mechanism, the method disentangles vision-sensitive tokens from high-entropy reasoning tokens and applies targeted feedback to each group. By integrating latent space optimization with anchor regularization, it achieves fine-grained, token-level coordination between perception and reasoning. Experimental results demonstrate that the proposed framework improves macro-average accuracy on models such as Qwen2.5-VL by 2.57% and 1.51%, respectively, compared to standard Chain-of-Thought (CoT) baselines, significantly outperforming existing methods.
📝 Abstract
Latent test-time scaling improves reasoning by refining hidden states during inference, but existing methods typically apply a single scalar reward to all editable latent tokens. For multimodal large language models, this global update ignores that generated tokens play different roles: some are sensitive to visual evidence, while others correspond to uncertain reasoning decisions. We present Token-Disentangled Latent Test-Time Scaling, an inference-time framework that makes latent refinement token-role-aware. Starting from an initial generated trajectory, we optimize a short hidden-state prefix while routing perception-side visual feedback to image-sensitive tokens and reasoning feedback to high-entropy tokens. Tokens selected by neither route are constrained by an anchor regularizer. Across both perception and reasoning benchmarks on Qwen2.5-VL-7B and InternVL3.5-8B, our method lifts macro accuracy over CoT by +2.57 and +1.51 respectively, and outperforms strong output-space test-time scaling baselines under matched decoded-candidate budgets. Code is available at https://github.com/Qwen-Applications/TD-LTTS.