Adaptive Sample Scheduling for Direct Preference Optimization

📅 2025-06-08
🏛️ arXiv.org
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
Static preference data in Direct Preference Optimization (DPO) mismatches the dynamically evolving model state during training, hindering optimization efficiency. Method: This work proposes the first adaptive sample scheduling paradigm tailored to LLM training—without altering DPO’s core algorithm. It constructs lightweight, online, batch-level sampling policies using multidimensional learning feedback signals: model output entropy, reward margin difference, and gradient sensitivity. Contribution/Results: (1) It formally defines and addresses the adaptive scheduling problem in DPO; (2) achieves an average 2.1% win-rate improvement across multiple alignment benchmarks, outperforming active learning and response-pair filtering baselines; (3) incurs negligible overhead (<0.5% latency increase), ensuring strong practicality and scalability.

Technology Category

Search and Optimization: Sampling/Simulation-based SearchMachine Learning: Active LearningPlanning, Routing, and Scheduling: Planning with Language Models

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
📝 Abstract
Direct Preference Optimization (DPO) has emerged as an effective approach for aligning large language models (LLMs) with human preferences. However, its performance is highly dependent on the quality of the underlying human preference data. To address this bottleneck, prior work has explored various data selection strategies, but these methods often overlook the impact of the evolving states of the language model during the DPO process. %including active querying, response pair selection, and data pre-selection. In this paper, we introduce a novel problem: Sample Scheduling for DPO, which aims to dynamically and adaptively schedule training samples based on the model's evolving states throughout preference optimization. To solve this problem, we propose SamS, an efficient and effective algorithm that adaptively selects samples in each training batch based on the LLM's learning feedback to maximize the potential generalization performance. Notably, without modifying the core DPO algorithm, simply integrating SamS significantly improves performance across tasks, with minimal additional computational overhead. This work points to a promising new direction for improving LLM alignment through more effective utilization of fixed preference datasets.
Problem

Research questions and friction points this paper is trying to address.

Optimizing sample selection for DPO using model feedback
Addressing performance dependency on human preference data quality
Dynamically scheduling training samples based on model states
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptively schedules training samples based on model states
Uses learning feedback to select samples for generalization
Enhances DPO without modifying its core algorithm
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.