planner-grounded fine-tuning

Designs and implements fine-tuning pipelines that incorporate offline planner-derived preference pairs and planner integration techniques, using Direct Preference Optimization (DPO) together with parameter-efficient LoRA updates so the model is optimized from planner signals without making online planner calls during training. Builds the optimization and inference machinery—LoRA+DPO recipes, training schedules, and interfaces—that enable inference-time planner-guided repair and allow the tuned model to be controlled or corrected by a planner at runtime.

planner-groundedfine-tuning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.01
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$184K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Pre-DPO: Improving Data Utilization in Direct Preference Optimization Using a Guiding Reference Model

Apr 22, 2025
JP
Junshu Pan
🏛️ Zhejiang University | Westlake University | Independent Researcher

In DPO, initializing the policy and reference models identically leads to inefficient data utilization and suboptimal performance; conversely, SimPO’s elimination of the reference model sacrifices robustness and risks catastrophic forgetting. To address this trade-off, we propose the Guided Reference Model (GRM), the first framework to formally characterize the reference model as a *dynamic sample-weighting mechanism* over preference pairs. GRM replaces the fixed reference model with a lightweight, learnable module that performs sample-level adaptive weighting—without introducing auxiliary reward models, external data, or additional parameters. This design seamlessly integrates DPO’s data efficiency with SimPO’s robustness. Empirical evaluation shows that GRM significantly outperforms standard DPO and SimPO on AlpacaEval 2.0 and Arena-Hard v0.1, delivering zero-cost improvements in reasoning alignment while incurring no increase in inference latency or model parameter count.

Addressing performance ceiling caused by identical policy and reference modelsEnhancing training robustness by leveraging a guiding reference modelImproving data utilization in Direct Preference Optimization (DPO)

Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks

Apr 23, 2024
AS
Amir Saeidi
🏛️ Arizona State University

This study systematically evaluates the effectiveness of Direct Preference Optimization (DPO) and its variants for aligning large language models (LLMs) with human preferences. We conduct a quantitative analysis across 13 multidimensional benchmarks—including MT-Bench, Big Bench, and the Open LLM Leaderboard—to assess DPO’s performance, dependence on supervised fine-tuning (SFT), the role of instruction tuning, and sensitivity to dataset scale. Key contributions: (1) DPO achieves 95% of full-dataset alignment performance using only 20% of preference data, demonstrating strong few-shot efficiency; (2) instruction tuning substantially improves factual consistency (+12.5%) and mathematical reasoning (+8.2%), though gains diminish on complex reasoning tasks; (3) we provide the first empirical evidence that DPO consistently enhances dialogue and question-answering performance (+3–7 percentage points). These findings establish theoretical foundations and practical guidelines for resource-efficient, high-fidelity LLM alignment.

Assess impact of training set size on performanceEvaluate DPO and variants for LLM alignmentTest alignment across diverse benchmark tasks

A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Oct 21, 2024
WX
Wenyi Xiao
🏛️ Zhejiang University | Nanyang Technological University | Alibaba Group

To address the high computational cost and reliance on reinforcement learning in RLHF-based alignment of large language models (LLMs), this paper presents a systematic survey of Direct Preference Optimization (DPO)—a reinforcement-learning-free alignment paradigm grounded solely in preference data. We introduce the first multidimensional taxonomy of DPO, unifying its theoretical foundations, algorithmic variants, benchmark datasets, and application domains. Through rigorous analysis grounded in Bradley–Terry modeling, loss function characterization, and data quality assessment, we empirically synthesize over 120 works to identify DPO’s convergence conditions, data sensitivity patterns, and scenario-specific adaptation strategies. Crucially, we uncover its fundamental theoretical limitations, training biases, and generalization bottlenecks for the first time. Finally, we propose three key future directions: scalability enhancement, robustness improvement, and multimodal extension—providing a principled methodological foundation for efficient, stable human preference alignment.

Aligning LLMs with human preferences efficientlyExploring future directions for model alignmentReviewing DPO's theories, variants, and limitations

This work addresses a key limitation of existing Direct Preference Optimization (DPO) methods, which rely solely on pairwise preference signals and neglect the quantitative differences in response quality, leading to ambiguous training signals and suboptimal optimization efficiency. To overcome this, we propose a novel preference optimization algorithm that, for the first time, incorporates explicit score gaps into the DPO framework. Our approach preserves the advantage of not requiring an explicit reward model while leveraging fine-grained relative quality information to enhance alignment. By designing a loss function that accounts for score differences, the method enjoys faster theoretical statistical convergence and demonstrates robustness to scoring noise. Extensive experiments show consistent and significant improvements over current DPO variants across multiple large language models and evaluation benchmarks, with stable performance gains even when provided with inaccurate scores.

Alignment ProblemDirect Preference OptimizationFoundation Models

MallowsPO: Fine-Tune Your LLM with Preference Dispersions

May 23, 2024
HC
Haoxian Chen
🏛️ Columbia University

This work addresses the limitation of Direct Preference Optimization (DPO) in modeling the diversity of human preferences. It introduces Mallows ranking theory—previously unexplored in LLM alignment—proposing a Mallows-based preference dispersion index that unifies and generalizes existing preference optimization frameworks. Methodologically, it parameterizes dispersion and reweights preference losses, enabling plug-and-play compatibility with mainstream offline preference optimization algorithms. Empirical evaluation on synthetic bandit, controllable generation, and dialogue tasks demonstrates significant performance gains; integrated as a lightweight plugin into Llama3-Instruct, it improves LC win rate by nearly 2%, confirming strong generalizability and interpretability. Core contributions are threefold: (1) a theoretically grounded, learnable metric for preference dispersion; (2) a unified generalization of DPO; and (3) an efficient, drop-in optimization enhancement for practical deployment.

Addresses DPO's inability to capture human preference diversity.Enhances DPO performance across various tasks and benchmarks.Introduces MallowsPO with a dispersion index for preference ranking.

Latest Papers

What's happening recently
View more

This work addresses the limitations of traditional reinforcement learning–based approaches for fine-tuning conversational agents, which often suffer from procedural complexity, high computational costs, and training instability. To overcome these challenges, the study proposes employing Direct Preference Optimization (DPO) as an alternative to reinforcement learning for efficiently fine-tuning large language models. The results demonstrate that DPO significantly simplifies the training pipeline and enhances training efficiency while maintaining strong generative performance, as evidenced by competitive scores on BLEU, ROUGE, and cosine similarity metrics. The research validates DPO’s effectiveness as a viable substitute for reinforcement learning in preference-based alignment and further identifies residual training instabilities under practical conditions, thereby providing an empirical foundation for future refinements.

Chatbot Fine-TuningDirect Preference OptimizationLarge Language Models

This study addresses the lack of systematic investigation into the interplay between supervised fine-tuning (SFT) and direct preference optimization (DPO), as well as parameter-efficient strategies, under conditions of small-scale language models and limited data. Using a GPT-2–sized decoder, the authors systematically compare training paradigms including SFT-only, DPO-only, and SFT followed by DPO, evaluating each with both full-parameter fine-tuning (FFT) and LoRA. Results demonstrate that FFT consistently and significantly outperforms LoRA, which fails to deliver practical speedups. Notably, when DPO’s preference construction aligns with the supervised objective, DPO alone—without SFT pretraining—achieves comparable performance, challenging the conventional assumption that SFT must precede preference-based alignment. The findings indicate that, in small-model settings, full-parameter SFT remains the dominant factor for achieving strong performance.

Direct Preference OptimizationEmpirical StudyParameterization

Existing preference optimization methods struggle to provide fine-grained feedback on multi-step solutions in complex reasoning tasks and face a trade-off between training stability and structured reasoning. To address this, this work proposes Hierarchical Preference Optimization (HiPO), a novel framework that introduces hierarchical structure into Direct Preference Optimization (DPO) for the first time. HiPO decomposes model responses into segments—such as query clarification, intermediate reasoning steps, and final answers—and computes DPO losses for each segment separately before fusing them with learned weights. This approach enables fine-grained alignment with human preferences while maintaining training efficiency. Experiments demonstrate that HiPO significantly outperforms DPO and other baselines across multiple 7B-scale models, achieving state-of-the-art performance on mathematical reasoning benchmarks and receiving higher evaluations from GPT-4.1 in terms of logical coherence, fluency, and consistency.

complex reasoningDirect Preference Optimizationfine-grained feedback

This work addresses a key limitation of conventional supervised fine-tuning (SFT), which merely imitates expert action sequences without discerning the relative quality of actions at the state level. To overcome this, the authors propose a lightweight offline policy optimization method that transforms expert trajectories into state-conditioned single-step action preferences. Leveraging a DPO-style objective, the approach contrasts expert actions against sampled negative actions—eliminating the need for environment interactions or an explicit reward model. The method introduces Policy-Preserving Augmentation (PPA) to maintain policy consistency and, for the first time, enables state-level action preference learning from expert-only data in a fully offline setting. Experiments demonstrate substantial improvements over SFT, boosting the accuracy of a 9B model on tau-bench from 21.7% to 41.4%, matching the performance of online GRPO.

expert trajectoriesimitation learningLLM agents

This work addresses the trade-off between the high computational cost of online reinforcement learning methods like GRPO and the limited cold-start performance of purely offline approaches such as DPO. The authors propose G2D, a three-stage framework that begins with minimal online GRPO pretraining to generate highly informative preference data, followed by constructing a static dataset through uncertainty calibration and difficulty-aware sampling, and finally fine-tuning offline using DPO. The study reveals that the performance gap between online and offline learning stems from insufficient data discriminability, and demonstrates that moderate pretraining mitigates overconfidence-induced information degradation. Experiments on Qwen2.5-7B and Llama-3.1-8B show that G2D achieves a 10.8% absolute improvement over GRPO (62.4% vs. 51.6%) while using only one-quarter of the computational budget, substantially enhancing both efficiency and performance.

Compute-Efficient RLData InformativenessOffline Preference Optimization

Hot Scholars

UT

Ufuk Topcu

The University of Texas at Austin
autonomycontrolsformal methodslearning
RS

Roni Stern

Software and Information Systems Engineering, Ben Gurion University of the Negev; Palo Alto Research
Artificial IntelligenceProgrammingComputer ScienceHeuristic Search
AA

Abhiroop Ajith

PhD Student, Worcester Polytechnic Institute
RoboticsDeep LearningAIManipulation
JB

Johannes Betz

Professor, Autonomous Vehicle Systems, Technical University of Munich (TUM)
Autonomous SystemsMotion PlaningControlRobots
CC

Constantinos Chamzas

Assistant Professor, Worcester Polytechnic Institute
RoboticsMotion PlanningPlanning Under UncertaintyLearning and Planning