BAS-OPD: Budget-Aware Selective On-Policy Self-Distillation for Fine-Grained Multimodal Perception

๐Ÿ“… 2026-09-22
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ๆœฌๆ–‡ๆๅ‡บBAS-OPDๆก†ๆžถ๏ผŒ้€š่ฟ‡้€‰ๆ‹ฉๆ€งๅœฐๅˆ†้…ๆœ‰้™็š„ๆ•™ๅธˆ็›‘็ฃ่ต„ๆบๆฅ่งฃๅ†ณๅคšๆจกๆ€ๅคงๆจกๅž‹ๅœจ็ป†็ฒ’ๅบฆ่ง†่ง‰ๆ„Ÿ็Ÿฅไธญ็š„้—ฎ้ข˜ใ€‚
๐Ÿ“ Abstract
Multimodal large language models (MLLMs) often struggle with fine-grained visual perception when processing complete images, as critical evidence may only appear in local regions. On-policy self-distillation (OPD) enables transferring privileged visual knowledge from informative views to full-image policies, but querying the teacher for every rollout introduces substantial supervision costs. In this work, we propose BAS-OPD, a budget-aware selective OPD framework that allocates teacher supervision under limited query budgets. Instead of querying all rollouts, BAS-OPD selects informative samples while maintaining full-batch student generation. We explore random, uncertainty-based, and learned utility-based selection strategies, where the learned selector estimates query value from detached rollout statistics and online utility signals derived from student--teacher agreement and teacher confidence without additional student forward passes. BAS-OPD only changes training-time supervision allocation and preserves single-pass full-image inference. Experiments on fine-grained multimodal perception benchmarks demonstrate that BAS-OPD achieves strong performance while substantially reducing teacher supervision costs, highlighting the effectiveness of selective OPD under constrained budgets.
Problem

Research questions and friction points this paper is trying to address.

Multimodal large language models
Fine-grained visual perception
On-policy self-distillation
Supervision costs
Query budget
Innovation

Methods, ideas, or system contributions that make the work stand out.

Budget-Aware
Selective On-Policy Self-Distillation
Multimodal Perception
Limited Query Budgets
Learned Selector
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
Z
Zihan Chen
HFIPS, Chinese Academy of Sciences; University of Science and Technology of China
H
Hengguang Zhou
University of California, Los Angeles
Y
Yuan Kang
HFIPS, Chinese Academy of Sciences; University of Science and Technology of China
Yiming Zhang
Yiming Zhang
University of Science and Technology of China
Computer Vision
W
Wenhui Fang
HFIPS, Chinese Academy of Sciences; University of Science and Technology of China
Z
Zenghui Ding
HFIPS, Chinese Academy of Sciences
Yining Sun
Yining Sun
Johns Hopkins University
Computer Vision
Cho-Jui Hsieh
Cho-Jui Hsieh
University of California, Los Angeles
Machine LearningOptimization