DEEPO: Dual-Entropy Enhanced Policy Optimization for Hallucination in MLLMs

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the hallucination problem in reinforcement learning for multimodal large language models, where high-entropy queries induce advantage collapse and the vanishing of confident error gradients. To mitigate this, we propose a dual-entropy enhancement strategy optimization framework featuring a novel two-stage mechanism. Specifically, semantic entropy triggers expert prefix injection to restore variance, while Rényi divergence preconditioning overcomes logarithmic saturation, thereby repairing the correction chain from reward signals to parameter updates. This approach significantly suppresses hallucinations while preserving generation fidelity. Empirical evaluations demonstrate a 4.0% improvement on the VideoMMMU benchmark, with statistically significant interaction effects and stable training dynamics throughout the optimization process.
📝 Abstract
Reinforcement learning (RL) is widely used to sharpen reasoning in multimodal large language models (MLLMs), yet its effect on hallucination is uneven. We trace this to two weak points in the \emph{correction chain} from reward to parameter update. At the rollout level, hard queries---those with high semantic entropy---frequently produce unanimously wrong sample groups, collapsing the group-relative advantage to zero exactly where hallucination risk is highest. At the optimization level, confident-but-wrong tokens are gradient-invisible: a categorical policy's expected score-gradient norm vanishes as its distribution sharpens, so the predictions that most need correction receive the weakest updates. We propose Dual-Entropy Enhanced Policy Optimization (DEEPO), a dual-stage enhancement combining signal variance regularization with gradient preconditioning: semantic-entropy-triggered expert prefixes inject grounded continuations on high-uncertainty queries, providing direct supervision and restoring advantage variance, while advantage-sign-aware Renyi preconditioning counteracts logit-level saturation so correction reaches confident errors in the operational confidence regime. Both branches improve over GRPO individually; their interaction is statistically significant on VideoMMMU---the most complex long-horizon task in our evaluation suite (+4.0$, 95\% CI [1.1, 6.9])---and additive elsewhere. DEEPO reduces hallucination while preserving accuracy and training stability.
Problem

Research questions and friction points this paper is trying to address.

Hallucination
Multimodal Large Language Models
Reinforcement Learning
Policy Optimization
Semantic Entropy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Policy Optimization
Hallucination Mitigation
Semantic Entropy
Gradient Preconditioning
Multimodal Large Language Models
🔎 Similar Papers