ChemOPD: Multi-Teacher On-Policy Distillation for Multi-Task Chemical Reasoning

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of integrating multi-task chemical reasoning capabilities into unified models by proposing a collaborative optimization framework based on expertized organization and guided integration. Methodologically, overlapping expert groups are constructed by combining supervised fine-tuning gradient estimation with constrained mixed-integer programming (MIP). Furthermore, an anchor-residual online distillation objective is designed to progressively introduce expert supervision atop a generalist teacher, thereby decoupling specialization from capability integration while enabling their synergy. Experimental results demonstrate that the proposed method yields significant improvements across multiple metrics on ChemCoTBench, outperforming single-expert distillation baselines and more fully unlocking the performance gains offered by multi-teacher settings.
📝 Abstract
Large language models are increasingly expected to support diverse chemical reasoning capabilities within a unified model. One approach is to develop specialized capabilities separately and consolidate them through multi-teacher on-policy distillation, but this raises two questions: how should specialization be organized, and how should specialist guidance be integrated? We introduce ChemOPD, which addresses both. We estimate task affinities from supervised fine-tuning gradients and solve a constrained mixed-integer program(MIP) to construct partially overlapping specialist groups. During distillation, we retain a generalist teacher trained on all tasks so that specialist guidance supplements rather than replaces its supervision. Our anchor-residual objective gradually increases the routed specialist's contribution on student-generated responses. On ChemCoTBench, affinity-guided specialization produces task-dependent gains over the generalist teacher and improves several capabilities beyond semantic task grouping. Yet stronger teacher-side performance does not automatically yield stronger students: with the same specialists and routes, anchor-residual OPD improves most reported metrics over specialist-only distillation and realizes a larger share of the available teacher gains. These results highlight specialization and capability integration as connected but distinct design problems in chemical reasoning.
Problem

Research questions and friction points this paper is trying to address.

chemical reasoning
multi-teacher distillation
task specialization
capability integration
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Distillation
Task Affinity
Mixed-Integer Programming
Anchor-Residual Objective
Chemical Reasoning
🔎 Similar Papers
2024-09-20Science China Information SciencesCitations: 0