🤖 AI Summary
This study addresses the limitations of offline fine-tuning, specifically that supervised fine-tuning (SFT) discards graded feedback and direct preference optimization (DPO) relies on paired data. To overcome these challenges, we propose a goal-conditioned supervised learning framework that translates feedback signals into explicit natural language objectives. Furthermore, we introduce a novel quality-threshold-based goal formulation that transcends the limitation of merely imitating high-quality subsets, enabling effective behavioral optimization of large language models through pure supervised learning. Experimental results demonstrate that our framework significantly outperforms existing baselines on tasks such as non-toxic generation, offering notable advantages in efficiency, scalability, and reduced data requirements.
📝 Abstract
Large language models often require fine-tuning to better align their behavior with user intent at deployment. Existing approaches are commonly divided into online and offline paradigms. Online methods, such as RL-based alignment, can directly optimize outcome quality but typically rely on external reward models and iterative rollouts, making them costly and difficult to deploy in many cases. Offline methods are more efficient, but prevailing approaches such as supervised fine-tuning (SFT) and direct preference optimization (DPO) remain limited: SFT typically collapses graded feedback into binary supervision, while DPO depends on paired preference data that is often unavailable or expensive to construct. In this paper, we propose goal-conditioned supervised learning (GCSL) as an offline fine-tuning framework for LLMs. Our core idea is to treat feedback signals directly as an explicit goal and train the model, purely through supervised learning, to generate responses that achieve that goal. To better exploit graded feedback, we further introduce a novel goal formulation that defines learning as consistently pursuing outcomes above a target quality threshold, rather than imitating samples from a selected high-quality subset. This design mitigates the bounded-learning effect by learning transferable patterns of meeting or exceeding quality thresholds. We also propose natural-language goal representations to further connect these patterns to the LLM's pretrained knowledge and generalization capabilities. We evaluate our method on three tasks: non-toxic generation, code generation, and LLM for recommendation. Results show that our approach consistently outperforms standard offline fine-tuning baselines while retaining the efficiency, scalability, and simple data requirements of supervised learning.