DRT: Dense Reasoning Trace for Efficient and Grounded Multimodal Reasoning

📅 2026-09-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决多模态大语言模型中信息稀释和视觉基础弱的问题,提出Dense Reasoning Trace方法,通过紧凑结构化追踪表达推理过程,提高令牌效率和逻辑基础。
📝 Abstract
Despite the remarkable progress in Multimodal Large Language Models (MLLMs), prevailing Chain-of-Thought (CoT) paradigms remain confined to the natural-language expression space. Consequently, they inherently incur excessive linguistic overhead, leading to information dilution and weak visual grounding. To address this challenge, we propose Dense Reasoning Trace (DRT), a paradigm that departs from natural-language-centered CoT by expressing reasoning as compact structured traces, which include concise intermediate states with symbolic connectors and disentangle visual observations from logical deductions. First, we introduce the Dense Trace Initialization to internalize the DRT reasoning mode into the model, substantially improving token efficiency while preserving visual evidence. To further enable the model to faithfully capture the logical relations within traces, we propose the Trace-Grounded Reinforcement Learning framework, which builds reference traces through a tri-perspective verification pipeline and employs Trace-Grounded GRPO with structured rewards, encouraging the model to generate concise DRT-style traces with reduced hallucination and stronger logical grounding. Extensive experiments on challenging reasoning benchmarks show that DRT achieves 5.5$\times$ token efficiency improvement while improving 1.3 accuracy points over the Qwen3-VL baseline. These findings suggest that complex multimodal reasoning may not require verbose natural-language traces, opening a more efficient path for next-generation MLLMs. Our code and data are available at: https://github.com/HIT-leaderone/DRT
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
Chain-of-Thought
visual grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dense Reasoning Trace
Dense Trace Initialization
Trace-Grounded Reinforcement Learning
💼 Related Jobs
No related jobs found.
W
Wan Xu
Harbin Institute of Technology
Y
Yuanfan Guo
FIS, ByteDance Inc
K
Kevin Han
Facebook
L
LaLa Chen
University of California, Irvine
Wangmeng Zuo
Wangmeng Zuo
School of Computer Science and Technology, Harbin Institute of Technology
Computer VisionImage ProcessingGenerative AIDeep LearningBiometrics