Score
Designs and builds pipelines that identify divergence points between rollout trajectories by matching states, construct pairwise preference comparisons from prefixes that share the same context, filter those pairs using action‑correctness annotations, and use the resulting context‑matched preference data to train preference‑based optimizers for sequential decision models under the same inference context.
This work addresses the challenges faced by multi-turn tool-using agents in long-horizon tasks, where coordinating tool sequences, tracking states, and enforcing strategic constraints are difficult to unify. Existing approaches suffer from a disconnect between reasoning and learning, leading to suboptimal tool selection and preference learning vulnerable to prompt misalignment. To overcome these limitations, the authors propose ToolGraph, a novel framework that integrates tool graph topology with divergence-point localization. By leveraging state matching and prefix alignment, ToolGraph identifies trajectory divergences and employs action correctness filtering to construct high-quality preference pairs, enabling context-consistent Direct Preference Optimization (DPO). Evaluated on 375 tau2-bench tasks, ToolGraph improves the weighted average reward from 0.304 to 0.338 (+11.2%), and further to 0.355 (+16.8%) when combined with DPO, substantially outperforming baselines—particularly in aviation and retail scenarios.
该研究通过引入FlowCPO解决了偏好对齐方法间关系不明确的问题,采用离线前向KL目标利用正负样本,无需在线采样。
This work addresses the problem of inferring an unknown preference pre-order defined over regular languages—representing temporal objectives—from users’ pairwise comparisons of finite-length trajectory sequences. To formalize such preferences, we introduce Preference Deterministic Finite Automata (PDFAs), the first automaton model encoding temporal objective preferences via partial-order-labeled transitions on deterministic finite automata. Theoretically, we prove that learning a minimal PDFA is NP-complete and propose a provably correct learning algorithm based on characteristic samples, with polynomial query complexity. Experimentally, our method accurately recovers user preferences in realistic robot motion planning scenarios. This work establishes a foundational bridge between formal language theory and preference learning, delivering both theoretical guarantees and empirical validation for preference inference over structured temporal specifications.
To address insufficient fine-grained preference alignment and weak error discrimination in large language model (LLM) tool invocation, this paper proposes the first token-level preference learning framework. Methodologically, we introduce a novel reverse data construction strategy, design a Token-level Preference Sampling (TPS) mechanism to capture local decision preferences during tool invocation, and propose an Error-guided Scoring Mechanism (ESM) for quantitative bias identification and correction. Through multi-turn tool-interaction fine-tuning, our framework achieves significant improvements in tool-call accuracy and robustness across three major benchmarks, demonstrating strong generalization across diverse models and datasets. The core contribution lies in advancing preference modeling from sentence-level to token-level—enabling precise behavioral modeling and fine-grained optimization of tool usage.
To address the high computational cost and reliance on reinforcement learning in RLHF-based alignment of large language models (LLMs), this paper presents a systematic survey of Direct Preference Optimization (DPO)—a reinforcement-learning-free alignment paradigm grounded solely in preference data. We introduce the first multidimensional taxonomy of DPO, unifying its theoretical foundations, algorithmic variants, benchmark datasets, and application domains. Through rigorous analysis grounded in Bradley–Terry modeling, loss function characterization, and data quality assessment, we empirically synthesize over 120 works to identify DPO’s convergence conditions, data sensitivity patterns, and scenario-specific adaptation strategies. Crucially, we uncover its fundamental theoretical limitations, training biases, and generalization bottlenecks for the first time. Finally, we propose three key future directions: scalability enhancement, robustness improvement, and multimodal extension—providing a principled methodological foundation for efficient, stable human preference alignment.
This work addresses the misalignment between the training objective of conventional conditional generative models and the downstream decision-making goal, which leads to significant errors in decision-sensitive regions that adversely affect optimal solutions. The authors propose Decision-Weighted Flow Matching (DW-FM), a framework that preserves the simplicity of flow matching while reweighting the velocity regression target using decision-sensitive information at the endpoints, thereby aligning training with downstream regret. By integrating loss-induced decision discrepancy with optimal transport theory, the paper establishes the first formal connection between pathwise velocity errors and decision regret, yielding a practical weighted objective with provable regret guarantees. Empirical results on synthetic portfolio allocation, semi-realistic financial, and traffic CVaR optimization tasks demonstrate that DW-FM substantially outperforms standard baselines and effectively reduces decision regret.
This work addresses the challenge of efficiently selecting the most informative pairwise comparisons under limited annotation budgets to improve alignment in preference-based large language model post-training. Framing comparison selection as a sampling design problem within the Direct Preference Optimization (DPO) framework, this study establishes the first theoretical connection between comparison pair sampling and policy suboptimality, deriving matching upper and lower bounds. Building on this analysis, the authors propose an explicit sampling criterion based on the Fisher information matrix to guide data acquisition. Experimental results demonstrate that the proposed method significantly outperforms existing heuristic strategies on both synthetic benchmarks and real-world language model post-training tasks, achieving substantially higher sample efficiency.
This study addresses the shallow shortcut bias in web process reward models (PRMs) caused by the scarcity of contrastive samples. To mitigate this, we propose SURFPRM, a framework that constructs an interactive element graph to systematically synthesize negative actions across spatial, temporal, and spatiotemporal dimensions. By integrating multi-strategy sampling, SURFPRM generates high-quality preference data through environment-grounded negative action proposals. Experimental results demonstrate that the framework increases the proportion of grounded minimal contrastive pairs from 24.19% to 74.60%. Furthermore, it surpasses existing baselines across multiple benchmarks, achieving performance comparable to proprietary large language models, and improves the success rate of the GPT-4o series on complex web tasks by over 12%.
This study addresses the challenges of local errors caused by distribution shift and reward scarcity in real-world scenarios during the deployment of flow-matching robotic policies. To this end, it proposes a test-time guidance framework based on QGF sampling. The method leverages human intervention data to train a preference model and innovatively introduces gradient upper bounds alongside zero-gradient losses to regularize the direction of preference gradients, thereby optimizing pretrained policies without updating the base model or requiring environmental rewards. Evaluated across four real-world precision insertion tasks, the proposed approach achieves an average success rate of 90.5%, representing a substantial improvement over the 69% attained by frozen policies. These results validate the effectiveness of utilizing intervention data as localized supervision for enhancing policy performance at test time.
Traditional alignment of large language models relies on handcrafted prompts or costly preference-based fine-tuning, which is inefficient and lacks interpretability. This work proposes Spec Learning, a novel framework that, for the first time, automatically generates human-readable and transparent natural language specifications from a small set of user preference pairs and brief instructions. During inference, these specifications conditionally guide model behavior without requiring any parameter updates. By circumventing black-box fine-tuning, Spec Learning outperforms Direct Preference Optimization (DPO) on preference-intensive, domain-specific datasets while providing interpretable behavioral rules. This approach significantly enhances both the efficiency and explainability of the alignment process.