🤖 AI Summary
This study addresses the severe performance degradation of multimodal large language models under extremely low visual token budgets. To this end, it proposes the LT-OPD framework, which introduces an online policy self-distillation mechanism enabling the compressed model to receive distributional supervision from a full-token teacher based on its own generation trajectories. Additionally, a progressive token budget curriculum learning strategy is designed to stabilize training under extreme compression conditions. Remarkably, when retaining only 5% of visual tokens, the proposed approach recovers 82.3% of the average performance relative to the full-token baseline while reducing KV cache and prefilling computation by over 85%, achieving an exceptional balance between efficient inference and model capability.
📝 Abstract
Visual token reduction is an effective way to accelerate multimodal large language models (MLLMs), but performance deteriorates rapidly under extremely low token budgets. Existing work has explored both visual-token selection and training-based adaptation to reduced visual inputs. We take a step further by asking how a heavily compressed MLLM should learn from the states induced by its own generations. This setting naturally calls for on-policy self-distillation: a heavily compressed model is supervised on the states induced by its own generations, while its full-token counterpart serves as an information-rich teacher. Based on this insight, we propose LT-OPD, a training framework for extreme visual-token reduction. The student rolls out responses with only a small fraction of visual tokens, and a frozen full-token copy of the same MLLM provides distributional supervision along these student-generated trajectories. To stabilize on-policy learning when visual evidence is severely limited, we further introduce a budget-level curriculum that progressively decreases the token budget during training. Across nine benchmarks on Qwen3.5-4B, LT-OPD raises average retained performance under 5% visual-token retention from 68.6% to 82.3%, outperforming training-free, training-based, and reinforcement-learning baselines at the same budget. The gains transfer consistently to Qwen3.5-9B, GLM-4.6V-9B, and LLaVA-OV-1.5-4B. LT-OPD also reduces KV-cache usage by 85.2% and prefill FLOPs by 85.4% without additional inference overhead, demonstrating that on-policy learning can substantially recover capabilities lost to extreme visual-token reduction.