🤖 AI Summary
This study addresses the imbalance between acquiring new capabilities and preserving general abilities during supervised fine-tuning (SFT) of offline agents. To this end, we propose a Privilege-Guided SFT method. The core innovation lies in introducing trajectory-turn-level information gain as an adaptive supervision signal for the first time, which dynamically modulates optimization intensity. This approach overcomes the limitations of conventional global constraints and KL-divergence penalties in preventing capability degradation. Experimental results demonstrate that the proposed method significantly mitigates distribution shift and broad capability deterioration, achieving a superior balance between new and general capabilities at the cost of only marginal performance loss on target tasks.
📝 Abstract
Supervised fine-tuning (SFT) on offline agent trajectories is the standard approach for training specialized tool-using agents, but forcing models to imitate reasoning and actions token by token may harm other capabilities (e.g., general reasoning, tool calling, code generation) of the base model. In this work, we focus on studying \emph{how to better balance the trade-off between acquiring new capabilities and preserving existing ones during agent trace SFT}. By comparing several baselines in our setup, standard SFT improves the target benchmark while lowering several non-target benchmark scores; meanwhile, simply constraining distributional drift using KL penalty or limiting the update magnitude did not avoid this regression trend. Motivated by recent token-wise adaptive learning objectives, this work proposes \textbf{Privilege-Guided SFT (PG-SFT)} to leverage turn-level information gain of agent trajectories as an indicator to adjust supervision strength. PG-SFT yields a more favorable observed trade-off on the evaluated benchmarks, substantially reducing distributional drift and broad capability degradation at the cost of slight degradation in target-task performance. Our findings suggest that balancing the acquisition--retention trade-off depends not only on whether the model is anchored to its base behavior, but also on where and how strongly supervision should depart from that behavior.}