divergence-point preference learning

Designs and builds pipelines that identify divergence points between rollout trajectories by matching states, construct pairwise preference comparisons from prefixes that share the same context, filter those pairs using action‑correctness annotations, and use the resulting context‑matched preference data to train preference‑based optimizers for sequential decision models under the same inference context.

divergence-pointpreferencelearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenges faced by multi-turn tool-using agents in long-horizon tasks, where coordinating tool sequences, tracking states, and enforcing strategic constraints are difficult to unify. Existing approaches suffer from a disconnect between reasoning and learning, leading to suboptimal tool selection and preference learning vulnerable to prompt misalignment. To overcome these limitations, the authors propose ToolGraph, a novel framework that integrates tool graph topology with divergence-point localization. By leveraging state matching and prefix alignment, ToolGraph identifies trajectory divergences and employs action correctness filtering to construct high-quality preference pairs, enabling context-consistent Direct Preference Optimization (DPO). Evaluated on 375 tau2-bench tasks, ToolGraph improves the weighted average reward from 0.304 to 0.338 (+11.2%), and further to 0.355 (+16.8%) when combined with DPO, substantially outperforming baselines—particularly in aviation and retail scenarios.

dialogue state trackingmulti-turn tool-usepreference learning

Automata Learning of Preferences over Temporal Logic Formulas from Pairwise Comparisons

May 23, 2025
HR
Hazhar Rahmani
🏛️ Missouri State University | University of Florida

This work addresses the problem of inferring an unknown preference pre-order defined over regular languages—representing temporal objectives—from users’ pairwise comparisons of finite-length trajectory sequences. To formalize such preferences, we introduce Preference Deterministic Finite Automata (PDFAs), the first automaton model encoding temporal objective preferences via partial-order-labeled transitions on deterministic finite automata. Theoretically, we prove that learning a minimal PDFA is NP-complete and propose a provably correct learning algorithm based on characteristic samples, with polynomial query complexity. Experimentally, our method accurately recovers user preferences in realistic robot motion planning scenarios. This work establishes a foundational bridge between formal language theory and preference learning, delivering both theoretical guarantees and empirical validation for preference inference over structured temporal specifications.

Develop algorithm to infer minimal PDFA from characteristic samplesLearn user preferences over temporal logic formulas from pairwise comparisonsModel preference relations using Preference Deterministic Finite Automata (PDFA)

TTPA: Token-level Tool-use Preference Alignment Training Framework with Fine-grained Evaluation

May 26, 2025
CH
Chengrui Huang
🏛️ University of Electronic Science and Technology of China | Shandong University

To address insufficient fine-grained preference alignment and weak error discrimination in large language model (LLM) tool invocation, this paper proposes the first token-level preference learning framework. Methodologically, we introduce a novel reverse data construction strategy, design a Token-level Preference Sampling (TPS) mechanism to capture local decision preferences during tool invocation, and propose an Error-guided Scoring Mechanism (ESM) for quantitative bias identification and correction. Through multi-turn tool-interaction fine-tuning, our framework achieves significant improvements in tool-call accuracy and robustness across three major benchmarks, demonstrating strong generalization across diverse models and datasets. The core contribution lies in advancing preference modeling from sentence-level to token-level—enabling precise behavioral modeling and fine-grained optimization of tool usage.

Align models with fine-grained tool-use preferencesImprove error discrimination in tool-learning methodsOptimize token-level tool call details in LLMs

A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

Oct 21, 2024
WX
Wenyi Xiao
🏛️ Zhejiang University | Nanyang Technological University | Alibaba Group

To address the high computational cost and reliance on reinforcement learning in RLHF-based alignment of large language models (LLMs), this paper presents a systematic survey of Direct Preference Optimization (DPO)—a reinforcement-learning-free alignment paradigm grounded solely in preference data. We introduce the first multidimensional taxonomy of DPO, unifying its theoretical foundations, algorithmic variants, benchmark datasets, and application domains. Through rigorous analysis grounded in Bradley–Terry modeling, loss function characterization, and data quality assessment, we empirically synthesize over 120 works to identify DPO’s convergence conditions, data sensitivity patterns, and scenario-specific adaptation strategies. Crucially, we uncover its fundamental theoretical limitations, training biases, and generalization bottlenecks for the first time. Finally, we propose three key future directions: scalability enhancement, robustness improvement, and multimodal extension—providing a principled methodological foundation for efficient, stable human preference alignment.

Aligning LLMs with human preferences efficientlyExploring future directions for model alignmentReviewing DPO's theories, variants, and limitations

Latest Papers

What's happening recently
View more

This work addresses the misalignment between the training objective of conventional conditional generative models and the downstream decision-making goal, which leads to significant errors in decision-sensitive regions that adversely affect optimal solutions. The authors propose Decision-Weighted Flow Matching (DW-FM), a framework that preserves the simplicity of flow matching while reweighting the velocity regression target using decision-sensitive information at the endpoints, thereby aligning training with downstream regret. By integrating loss-induced decision discrepancy with optimal transport theory, the paper establishes the first formal connection between pathwise velocity errors and decision regret, yielding a practical weighted objective with provable regret guarantees. Empirical results on synthetic portfolio allocation, semi-realistic financial, and traffic CVaR optimization tasks demonstrate that DW-FM substantially outperforms standard baselines and effectively reduces decision regret.

conditional generative modelsdecision regretobjective mismatch

This work addresses the challenge of efficiently selecting the most informative pairwise comparisons under limited annotation budgets to improve alignment in preference-based large language model post-training. Framing comparison selection as a sampling design problem within the Direct Preference Optimization (DPO) framework, this study establishes the first theoretical connection between comparison pair sampling and policy suboptimality, deriving matching upper and lower bounds. Building on this analysis, the authors propose an explicit sampling criterion based on the Fisher information matrix to guide data acquisition. Experimental results demonstrate that the proposed method significantly outperforms existing heuristic strategies on both synthetic benchmarks and real-world language model post-training tasks, achieving substantially higher sample efficiency.

comparison selectionlabeling budgetLLM alignment

This study addresses the shallow shortcut bias in web process reward models (PRMs) caused by the scarcity of contrastive samples. To mitigate this, we propose SURFPRM, a framework that constructs an interactive element graph to systematically synthesize negative actions across spatial, temporal, and spatiotemporal dimensions. By integrating multi-strategy sampling, SURFPRM generates high-quality preference data through environment-grounded negative action proposals. Experimental results demonstrate that the framework increases the proportion of grounded minimal contrastive pairs from 24.19% to 74.60%. Furthermore, it surpasses existing baselines across multiple benchmarks, achieving performance comparable to proprietary large language models, and improves the success rate of the GPT-4o series on complex web tasks by over 12%.

Grounded Minimal Contrastive PairsPreference Data SynthesisProcess Reward Models

This study addresses the challenges of local errors caused by distribution shift and reward scarcity in real-world scenarios during the deployment of flow-matching robotic policies. To this end, it proposes a test-time guidance framework based on QGF sampling. The method leverages human intervention data to train a preference model and innovatively introduces gradient upper bounds alongside zero-gradient losses to regularize the direction of preference gradients, thereby optimizing pretrained policies without updating the base model or requiring environmental rewards. Evaluated across four real-world precision insertion tasks, the proposed approach achieves an average success rate of 90.5%, representing a substantial improvement over the 69% attained by frozen policies. These results validate the effectiveness of utilizing intervention data as localized supervision for enhancing policy performance at test time.

distribution shiftflow-matching policyhuman interventions

Traditional alignment of large language models relies on handcrafted prompts or costly preference-based fine-tuning, which is inefficient and lacks interpretability. This work proposes Spec Learning, a novel framework that, for the first time, automatically generates human-readable and transparent natural language specifications from a small set of user preference pairs and brief instructions. During inference, these specifications conditionally guide model behavior without requiring any parameter updates. By circumventing black-box fine-tuning, Spec Learning outperforms Direct Preference Optimization (DPO) on preference-intensive, domain-specific datasets while providing interpretable behavioral rules. This approach significantly enhances both the efficiency and explainability of the alignment process.

inference-time steeringlarge language modelspreference alignment