SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high inference cost of vision-language models, where aggressive visual token pruning significantly degrades performance. To mitigate this, we propose SCOPD, a framework leveraging sparse-context online self-distillation that enables a student model to generate reasoning trajectories on pruned tokens under the supervision of a full-context teacher without additional inference overhead. Furthermore, by identifying a "representation-utilization gap," we introduce SCOPD+, which incorporates a selective distillation strategy to optimize vision-sensitive positions. Under an extreme setting retaining only 10% of tokens, SCOPD and SCOPD+ recover performance to 90.49% and 92.43%, respectively, substantially outperforming existing baselines.
📝 Abstract
Reasoning vision-language models (VLMs) process images and videos as long sequences of visual tokens, making inference expensive. Training-free token pruning reduces this cost, but aggressive compression can sharply degrade performance, often attributed to irreversible loss of task-relevant visual information. We show that this explanation is incomplete. In a fixed-context Pass@K analysis, repeated sampling from the same pruned visual representation recovers many examples missed by greedy decoding, indicating that useful visual evidence can remain accessible but be used unreliably. We call this the representation-utilization gap. Motivated by this observation, we introduce SCOPD, a sparse-context on-policy self-distillation framework in which a student generates reasoning trajectories from pruned visual tokens while a privileged full-context teacher supervises the same on-policy prefixes. SCOPD requires no ground-truth responses, architectural changes, or additional inference-time computation. We further introduce SCOPD+, which uses a small visual-budget intervention to identify visually sensitive response positions and selectively distill them. At 10% visual-token retention, the Vanilla model retains 86.37% of its unpruned performance across 13 benchmarks. SCOPD raises this to 90.49%, while SCOPD+ further improves it to 92.43%. Across token budgets, benchmarks, and pruning operators, our results show that efficient reasoning depends not only on which visual information survives pruning, but also on how reliably the model learns to use it.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Token Pruning
Representation-Utilization Gap
Efficient Inference
Visual Token Compression
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Distillation
Token Pruning
Vision-Language Models
Sparse-Context
On-Policy
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.