Lexicographic Multi-Objective On-Policy Distillation

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of multi-reward post-training, where the absence of explicit prioritization often degrades high-priority capabilities such as correctness. To overcome this, we propose LMOPD, a framework integrating reinforcement learning, multi-teacher distillation, and a Mixture-of-Experts (MoE) architecture. LMOPD establishes explicit priorities among expert policies through lexicographic routing and introduces a local log-policy projection mechanism to achieve optimal expert ensembling under asymmetric trade-offs. Evaluated on mathematical reasoning benchmarks, our approach preserves approximately 90% accuracy while yielding substantial reasoning gains, significantly outperforming existing baselines. This work effectively resolves the critical challenge in multi-objective optimization where high-priority capabilities are eroded by lower-priority objectives.
📝 Abstract
Reinforcement learning from verifiable rewards (RLVR) usually optimizes answer correctness, yet useful language-model behavior also requires high-quality reasoning and concise responses. Existing multi-reward post-training methods typically scalarize rewards or combine specialists without explicitly protecting a reward priority order. This is problematic when trade-offs are asymmetric: conciseness, for example, should not improve at the cost of correctness. We introduce Lexicographic Multi-Objective On-Policy Distillation (LMOPD), a multi-teacher method for integrating reward-specialized policies under explicit priorities. For each student rollout, LMOPD selects the specialist for the first objective whose gate detects a deficiency, then locally projects its centered log-policy correction to remove components that oppose higher-priority specialists. We evaluate 30B-A3B mixture-of-experts transformer models in two- and four-expert settings on three math benchmarks, measuring retained specialist gains. With two experts, LMOPD's point estimates fully retain the accuracy and reasoning-quality gains while acquiring $46.9\%$ of the conciseness gain. With four experts, it retains $\approx90\%$ of both the accuracy gain and reasoning-correctness gain, compared to only $\approx57\%$ by the next best evaluated baseline. Matched four-expertablations show that lexicographic routing outperforms random routing and that projection further strengthens both top-priority capabilities. Across both scales, LMOPD preserves the highest-priority capabilities more effectively than the existing baselines we evaluate, demonstrating the value of explicit priorities for specialist integration.
Problem

Research questions and friction points this paper is trying to address.

multi-objective reinforcement learning
reward prioritization
policy distillation
asymmetric trade-offs
RLVR
Innovation

Methods, ideas, or system contributions that make the work stand out.

Lexicographic Multi-Objective Optimization
On-Policy Distillation
Multi-Teacher Integration
Policy Projection
Reinforcement Learning from Verifiable Rewards
🔎 Similar Papers
2024-07-21arXiv.orgCitations: 1