Score
Designs and implements fine-tuning pipelines that incorporate offline planner-derived preference pairs and planner integration techniques, using Direct Preference Optimization (DPO) together with parameter-efficient LoRA updates so the model is optimized from planner signals without making online planner calls during training. Builds the optimization and inference machinery—LoRA+DPO recipes, training schedules, and interfaces—that enable inference-time planner-guided repair and allow the tuned model to be controlled or corrected by a planner at runtime.
This paper addresses the limitations of large language model (LLM)-based agents in long-horizon planning, dynamic interaction, and complex decision-making within intricate environments. Methodologically, it introduces the first systematic optimization survey, proposing a unified classification framework that dichotomizes optimization strategies into parameter-driven approaches (e.g., supervised fine-tuning, PPO, DPO) and parameter-agnostic techniques (e.g., prompt engineering, retrieval-augmented generation, reward shaping, trajectory construction). It further analyzes critical integrative aspects—such as hybrid optimization—and synthesizes evaluation benchmarks and representative applications. Contributions include: (1) a structured, comprehensive review encompassing over 100 works; (2) an open-source, standardized reference library hosted on GitHub; and (3) a clear articulation of open challenges and actionable research directions. Collectively, this work provides both theoretical foundations and a reproducible toolchain for efficient LLM agent optimization.
In DPO, initializing the policy and reference models identically leads to inefficient data utilization and suboptimal performance; conversely, SimPO’s elimination of the reference model sacrifices robustness and risks catastrophic forgetting. To address this trade-off, we propose the Guided Reference Model (GRM), the first framework to formally characterize the reference model as a *dynamic sample-weighting mechanism* over preference pairs. GRM replaces the fixed reference model with a lightweight, learnable module that performs sample-level adaptive weighting—without introducing auxiliary reward models, external data, or additional parameters. This design seamlessly integrates DPO’s data efficiency with SimPO’s robustness. Empirical evaluation shows that GRM significantly outperforms standard DPO and SimPO on AlpacaEval 2.0 and Arena-Hard v0.1, delivering zero-cost improvements in reasoning alignment while incurring no increase in inference latency or model parameter count.
This study systematically evaluates the effectiveness of Direct Preference Optimization (DPO) and its variants for aligning large language models (LLMs) with human preferences. We conduct a quantitative analysis across 13 multidimensional benchmarks—including MT-Bench, Big Bench, and the Open LLM Leaderboard—to assess DPO’s performance, dependence on supervised fine-tuning (SFT), the role of instruction tuning, and sensitivity to dataset scale. Key contributions: (1) DPO achieves 95% of full-dataset alignment performance using only 20% of preference data, demonstrating strong few-shot efficiency; (2) instruction tuning substantially improves factual consistency (+12.5%) and mathematical reasoning (+8.2%), though gains diminish on complex reasoning tasks; (3) we provide the first empirical evidence that DPO consistently enhances dialogue and question-answering performance (+3–7 percentage points). These findings establish theoretical foundations and practical guidelines for resource-efficient, high-fidelity LLM alignment.
To address the high computational cost and reliance on reinforcement learning in RLHF-based alignment of large language models (LLMs), this paper presents a systematic survey of Direct Preference Optimization (DPO)—a reinforcement-learning-free alignment paradigm grounded solely in preference data. We introduce the first multidimensional taxonomy of DPO, unifying its theoretical foundations, algorithmic variants, benchmark datasets, and application domains. Through rigorous analysis grounded in Bradley–Terry modeling, loss function characterization, and data quality assessment, we empirically synthesize over 120 works to identify DPO’s convergence conditions, data sensitivity patterns, and scenario-specific adaptation strategies. Crucially, we uncover its fundamental theoretical limitations, training biases, and generalization bottlenecks for the first time. Finally, we propose three key future directions: scalability enhancement, robustness improvement, and multimodal extension—providing a principled methodological foundation for efficient, stable human preference alignment.
This work addresses a key limitation of existing Direct Preference Optimization (DPO) methods, which rely solely on pairwise preference signals and neglect the quantitative differences in response quality, leading to ambiguous training signals and suboptimal optimization efficiency. To overcome this, we propose a novel preference optimization algorithm that, for the first time, incorporates explicit score gaps into the DPO framework. Our approach preserves the advantage of not requiring an explicit reward model while leveraging fine-grained relative quality information to enhance alignment. By designing a loss function that accounts for score differences, the method enjoys faster theoretical statistical convergence and demonstrates robustness to scoring noise. Extensive experiments show consistent and significant improvements over current DPO variants across multiple large language models and evaluation benchmarks, with stable performance gains even when provided with inaccurate scores.
This work addresses the limitation of Direct Preference Optimization (DPO) in modeling the diversity of human preferences. It introduces Mallows ranking theory—previously unexplored in LLM alignment—proposing a Mallows-based preference dispersion index that unifies and generalizes existing preference optimization frameworks. Methodologically, it parameterizes dispersion and reweights preference losses, enabling plug-and-play compatibility with mainstream offline preference optimization algorithms. Empirical evaluation on synthetic bandit, controllable generation, and dialogue tasks demonstrates significant performance gains; integrated as a lightweight plugin into Llama3-Instruct, it improves LC win rate by nearly 2%, confirming strong generalizability and interpretability. Core contributions are threefold: (1) a theoretically grounded, learnable metric for preference dispersion; (2) a unified generalization of DPO; and (3) an efficient, drop-in optimization enhancement for practical deployment.
This work addresses the limitations of traditional reinforcement learning–based approaches for fine-tuning conversational agents, which often suffer from procedural complexity, high computational costs, and training instability. To overcome these challenges, the study proposes employing Direct Preference Optimization (DPO) as an alternative to reinforcement learning for efficiently fine-tuning large language models. The results demonstrate that DPO significantly simplifies the training pipeline and enhances training efficiency while maintaining strong generative performance, as evidenced by competitive scores on BLEU, ROUGE, and cosine similarity metrics. The research validates DPO’s effectiveness as a viable substitute for reinforcement learning in preference-based alignment and further identifies residual training instabilities under practical conditions, thereby providing an empirical foundation for future refinements.
This study addresses the lack of systematic investigation into the interplay between supervised fine-tuning (SFT) and direct preference optimization (DPO), as well as parameter-efficient strategies, under conditions of small-scale language models and limited data. Using a GPT-2–sized decoder, the authors systematically compare training paradigms including SFT-only, DPO-only, and SFT followed by DPO, evaluating each with both full-parameter fine-tuning (FFT) and LoRA. Results demonstrate that FFT consistently and significantly outperforms LoRA, which fails to deliver practical speedups. Notably, when DPO’s preference construction aligns with the supervised objective, DPO alone—without SFT pretraining—achieves comparable performance, challenging the conventional assumption that SFT must precede preference-based alignment. The findings indicate that, in small-model settings, full-parameter SFT remains the dominant factor for achieving strong performance.
Existing preference optimization methods struggle to provide fine-grained feedback on multi-step solutions in complex reasoning tasks and face a trade-off between training stability and structured reasoning. To address this, this work proposes Hierarchical Preference Optimization (HiPO), a novel framework that introduces hierarchical structure into Direct Preference Optimization (DPO) for the first time. HiPO decomposes model responses into segments—such as query clarification, intermediate reasoning steps, and final answers—and computes DPO losses for each segment separately before fusing them with learned weights. This approach enables fine-grained alignment with human preferences while maintaining training efficiency. Experiments demonstrate that HiPO significantly outperforms DPO and other baselines across multiple 7B-scale models, achieving state-of-the-art performance on mathematical reasoning benchmarks and receiving higher evaluations from GPT-4.1 in terms of logical coherence, fluency, and consistency.
This work addresses a key limitation of conventional supervised fine-tuning (SFT), which merely imitates expert action sequences without discerning the relative quality of actions at the state level. To overcome this, the authors propose a lightweight offline policy optimization method that transforms expert trajectories into state-conditioned single-step action preferences. Leveraging a DPO-style objective, the approach contrasts expert actions against sampled negative actions—eliminating the need for environment interactions or an explicit reward model. The method introduces Policy-Preserving Augmentation (PPA) to maintain policy consistency and, for the first time, enables state-level action preference learning from expert-only data in a fully offline setting. Experiments demonstrate substantial improvements over SFT, boosting the accuracy of a 9B model on tau-bench from 21.7% to 41.4%, matching the performance of online GRPO.
This work addresses the trade-off between the high computational cost of online reinforcement learning methods like GRPO and the limited cold-start performance of purely offline approaches such as DPO. The authors propose G2D, a three-stage framework that begins with minimal online GRPO pretraining to generate highly informative preference data, followed by constructing a static dataset through uncertainty calibration and difficulty-aware sampling, and finally fine-tuning offline using DPO. The study reveals that the performance gap between online and offline learning stems from insufficient data discriminability, and demonstrates that moderate pretraining mitigates overconfidence-induced information degradation. Experiments on Qwen2.5-7B and Llama-3.1-8B show that G2D achieves a 10.8% absolute improvement over GRPO (62.4% vs. 51.6%) while using only one-quarter of the computational budget, substantially enhancing both efficiency and performance.