LEGO-OPD: Factorized Teacher Composition for Multimodal On-Policy Distillation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the deep entanglement of visual signals and linguistic priors in multimodal distillation, which hinders simultaneously enhancing perception and preserving reasoning. We propose LEGO-OPD, a framework that introduces a generalized Bayesian factorized teacher composition mechanism. By independently constructing language and localization experts to form a unified teacher distribution, it precisely decouples reasoning from perceptual supervision. Furthermore, we design a prefix-dependent adaptive calibration strategy that dynamically modulates the intensity of visual likelihood updates to prevent over- or under-supervision. Experiments on Qwen3-based multimodal online policy distillation demonstrate that our method significantly outperforms baselines across both multimodal and text-only tasks, effectively improving visual perception capabilities while fully preserving linguistic reasoning performance.
📝 Abstract
Multimodal on-policy distillation (OPD) aims to improve visual grounding while preserving the strong reasoning capabilities of language models. Recent multi-teacher approaches combine LLM and VLM teachers to provide complementary supervision. However, directly using a VLM's full predictive distribution entangles its visual grounding signal with its own language prior, preventing the grounding information from being transferred independently. Conversely, increasing the strength of visual supervision can improve perception but may overemphasize visual evidence and degrade language reasoning. To address this trade-off, we introduce LEGO-OPD, which selectively composes factors from a Language Expert and a Grounding expert into One teacher distribution for multimodal OPD. Under a generalized Bayesian formulation, the language expert provides a prior over candidate tokens, while the grounding expert contributes a visual likelihood that updates this prior, rather than transferring its complete predictive distribution. This factorized composition allows language reasoning and visual grounding to be controlled independently. We further introduce adaptive calibration to determine how strongly the visual likelihood should update the language prior at each decoding prefix. Specifically, LEGO-OPD uses the grounding expert's image-induced prediction shift as a prefix-dependent reference, preventing both insufficient and excessive visual supervision. Experiments with Qwen3 models show that LEGO-OPD consistently outperforms the evaluated single- and multi-teacher OPD baselines on both multimodal and text-only reasoning tasks. Moreover, it improves the initial student's visual perception while preserving text-only reasoning.
Problem

Research questions and friction points this paper is trying to address.

multimodal on-policy distillation
visual grounding
language reasoning
teacher composition
predictive distribution entanglement
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal On-Policy Distillation
Factorized Teacher Composition
Generalized Bayesian Formulation
Adaptive Calibration
Visual Grounding
🔎 Similar Papers
No similar papers found.
J
Jaeyun Shin
Korea Advanced Institute of Science and Technology (KAIST)
H
Hangeol Chang
Korea Advanced Institute of Science and Technology (KAIST)
Jong Chul Ye
Jong Chul Ye
Professor, Chung Moon Soul Chair, Graduate School of AI, KAIST
machine learningcomputational imagingmedical imagingsignal processingcompressed sensing