Beyond Entropy: Self-Diagnostic Multi-Role Token Optimization for Video Reasoning

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of ambiguous token-level credit assignment and visual exploration interference in reinforcement learning for video reasoning by proposing the DyCPO framework. Methodologically, it introduces a multi-role dependency metric to balance visual exploration with answer mining, and leverages self-diagnostic dynamic counterfactual signals to optimize token selection. The core innovation lies in establishing the first co-evolutionary mechanism between policy and optimization objectives, replacing static priors with dynamic synergistic evolution. Experimental results demonstrate that this approach significantly enhances performance on complex video reasoning benchmarks, establishing a robust new paradigm for token-level credit assignment.
📝 Abstract
Reinforcement learning with verifiable rewards has substantially advanced multimodal reasoning, yet it remains fundamentally limited by ambiguous token-level credit assignment. While high-entropy token heuristics encourage possibility exploration, naively extending them to video reasoning tends to induce lengthy reasoning, as the model becomes overly reliant on high-entropy visual activations. Alternative approaches that rely on counterfactual-based visual token localization for credit assignment also tend to over-prioritize visual exploration at the expense of decisive reasoning cues for answer derivation, thereby exacerbating the interference from spurious visual nuances. Moreover, these methods employ static counterfactual strategies that fail to co-evolve with the policy during training. In this paper, we introduce DyCPO, a co-evolutionary framework that jointly optimizes reliable token selection and adaptive counterfactual intervention. It constructs a multi-role dependence metric to balance visual exploration and answer-relevance mining in token-wise contrastive learning, while suppressing exploration-only filler tokens and spurious visual noise. Rather than relying on static counterfactual priors, DyCPO dynamically derives counterfactual signals from the model's own successful and failed rollouts, enabling self-diagnostic analysis and co-evolution of the optimization objective with the policy. Extensive experiments on complex video reasoning and general video understanding benchmarks demonstrate consistent performance improvements, establishing DyCPO as a robust token-level credit assignment paradigm for multimodal reinforcement learning.
Problem

Research questions and friction points this paper is trying to address.

video reasoning
credit assignment
reinforcement learning
token optimization
counterfactual intervention
Innovation

Methods, ideas, or system contributions that make the work stand out.

Token-level Credit Assignment
Dynamic Counterfactual Intervention
Multi-Role Dependence Metric
Co-evolutionary Optimization
Video Reasoning
🔎 Similar Papers
2024-02-20International Conference on Machine LearningCitations: 30