Score
Designs and implements algorithms that construct preference data by contrasting failed and successful trajectories, treating failures as negative examples and successes as positives. Builds contrastive preference models that adapt policy objectives or ranking functions so agents avoid failed behaviors, enabling preference learning from failures without requiring human labels.
In real-world settings, high-quality task-specific data is scarce, causing agents to overfit by excessively relying on limited expert demonstrations. Method: This paper proposes a Coevolutionary Agent Framework that introduces a collaborative evolution mechanism between a target agent and a failure agent. The failure agent actively generates “near-success but ultimately failing” hard negative samples, transforming structured failures into discriminative supervisory signals. Integrated with preference optimization, self-generated trajectory sampling, and hard negative mining, the framework enables interactive, inter-agent improvement. Results: It significantly outperforms supervised fine-tuning across multiple benchmarks, demonstrating markedly improved generalization. Core contribution: This work is the first to model controllable failure as a principled source of negative samples—breaking the traditional self-improvement paradigm’s exclusive reliance on positive trajectories—and thereby effectively mitigates overfitting while enhancing decision-boundary learning.
This work investigates the fundamental limitations of ordinal preference feedback (e.g., pairwise comparisons) for optimizing large language model outputs in complex human feedback tasks—such as deep research or travel planning. Using the first formal integration of social choice theory (particularly voting theory) into RLHF analysis, and combining it with reinforcement learning theory and preference learning generalization bounds, we rigorously prove that even under ideal conditions—infinite data, zero noise, and online preference acquisition—post-training based solely on ordinal feedback cannot guarantee convergence to an approximately optimal policy. We further disentangle distinct failure modes across reasoning-oriented settings versus instruction tuning, exposing inherent unreliability in ordinal feedback. Our core contribution is establishing a theoretical bottleneck for RLHF in complex reasoning tasks, and demonstrating that overcoming it necessitates incorporating absolute (cardinal) scoring mechanisms and designing novel algorithms grounded in richer feedback structures.
This study addresses the lack of supervision for boundary failure samples in offline preference optimization, where responses violate the original instruction yet satisfy semantically adjacent intents. To this end, we propose a bidirectional preference synthesis method that constructs forward and reverse paired data, introducing reverse preference pairs under "achieved prompts" so that identical responses are rejected in incorrect contexts while selected in correct ones. This explicitly models prompt-conditioned dependencies, refining supervision signals without modifying the DPO objective or training a reward model. Experiments demonstrate that our approach improves achieved-side ranking accuracy from 6.8% to 62.3%, significantly enhancing multilingual multi-turn instruction-following capabilities while maintaining consistent advantages across agent-based, tool-calling, and code generation tasks.
This study investigates the optimal timing for incorporating preference embeddings in artificial intelligence systems—whether during training or as a post-processing step. Drawing on information design theory, the work proposes a unified welfare framework that is agnostic to specific decision objectives, revealing how preference embedding induces a contraction in posterior means and thereby affects the value of information. By integrating convexity analysis of information value with models of cognitive constraints, the paper demonstrates that, in the absence of cognitive frictions, preference-agnostic training weakly dominates preference-embedded approaches. However, under human cognitive limitations, embedding preferences during training enhances performance through automatic threshold computation, offering theoretical justification for modular AI architectures.
This work addresses a key limitation of existing Direct Preference Optimization (DPO) methods, which rely solely on pairwise preference signals and neglect the quantitative differences in response quality, leading to ambiguous training signals and suboptimal optimization efficiency. To overcome this, we propose a novel preference optimization algorithm that, for the first time, incorporates explicit score gaps into the DPO framework. Our approach preserves the advantage of not requiring an explicit reward model while leveraging fine-grained relative quality information to enhance alignment. By designing a loss function that accounts for score differences, the method enjoys faster theoretical statistical convergence and demonstrates robustness to scoring noise. Extensive experiments show consistent and significant improvements over current DPO variants across multiple large language models and evaluation benchmarks, with stable performance gains even when provided with inaccurate scores.
This paper addresses key challenges in constrained reinforcement learning (CRL) for safety-critical control, preference learning, and large language model (LLM) alignment. First, for average-cost and finite-horizon constrained Markov decision processes (CMDPs), we propose ACPO and e-COP—algorithms achieving both theoretical optimality and state-of-the-art constraint satisfaction. Second, for preference-based learning, we introduce warmPref-PS and PSPL, the first methods integrating uncertainty quantification with posterior sampling, significantly reducing regret and enhancing robust policy identification under sparse or noisy preferences. Third, for multi-objective LLM alignment, we develop MOPO—a scalable, constraint-aware optimization framework supporting billion-parameter models. All methods are underpinned by rigorous theoretical guarantees—including convergence, constraint violation bounds, and regret analysis—and empirically validated across diverse real-world benchmarks in robotics, human-AI interaction, and LLM fine-tuning.
This work addresses the challenges of scarce and costly demonstration data in real-world settings and the error compounding in conventional imitation learning due to its reliance on the i.i.d. assumption during testing. To overcome these limitations, the paper introduces the “Master Your Own Expertise” (MYOE) framework, which integrates a Queryable Mixture-of-Preferences State-Space Model (QMoP-SSM) with a preference-based regret mechanism. MYOE estimates desired targets at each step and refines a neural control policy through self-imitation from limited demonstrations. By unifying reinforcement learning, imitation learning, state-space modeling, and preference reasoning, MYOE circumvents the heavy dependence of existing RLfD approaches on large-scale data and distributional consistency. Experimental results demonstrate that MYOE significantly outperforms state-of-the-art methods in robustness, adaptability, and out-of-distribution generalization.
Although contrastive critics effectively rank actions, their objective functions are prone to selecting out-of-distribution actions during expressive policy search due to embedding norm drift and cosine binding. This work is the first to systematically identify the root cause as a decoupling between ranking capability and value calibration, diagnosing the issue through support set decomposition and disentangling training from readout. Leveraging TD-Q baselines with Bellman-consistent training, OGBench navigation tasks, and simulator rollouts, the study reveals substantial one-step selection costs in PointMaze and Q* tasks, while controllers in AntMaze and HumanoidMaze exhibit self-correction. Notably, even after inference-time normalization, original embeddings retain weak ranking ability. The findings elucidate how training objectives and inference normalization jointly shape ranking performance.
Traditional offline evaluation relies solely on task success or failure, ignoring intermediate progress and resulting in identical assessments for a large proportion of samples, which undermines statistical efficiency and system discriminability. This work introduces trajectory-level temporal preferences into offline evaluation for the first time, proposing a preference-based trajectory evaluation method that directly compares agent trajectories through progress-aware preferences and return distributions over time. The approach establishes an evaluation framework integrating preference modeling, pairwise trajectory comparison, and joint temporal-progress analysis. Evaluated across diverse agents and interactive benchmarks, it reduces evaluation tie rates from approximately 75% to 35%, substantially improving discriminative power, ranking stability, and data efficiency. Furthermore, the study reveals that benchmark saturation may stem not from agent capabilities but from limitations inherent in conventional evaluation metrics themselves.
This work addresses the challenge of specifying and balancing optimization objectives for robots in complex environments, where manual design is difficult and existing active preference learning methods are constrained by fixed trajectory sets, limiting query informativeness and diversity. To overcome this, the authors propose a novel approach that jointly optimizes environment design and trajectory selection for the first time. By leveraging counterfactual reasoning, the method generates trajectory pairs that effectively reveal differences among candidate reward functions, thereby eliciting more informative user preferences. The framework integrates Bayesian reward belief sampling, learnable environment parameterization, and a strategic trajectory-pair generation policy to substantially enhance both the information content and diversity of queries. Experiments demonstrate that the proposed method outperforms existing approaches in reward accuracy and sample efficiency, and achieves higher user ratings in human subject studies.