reference-guided policy improvement

Designs and implements algorithms and training procedures that improve a learned policy by injecting external reference trajectories or solutions into rollout data and optimization, including mechanisms to detect and target low-performing rollout groups, convert blind exploration into guided improvement, and bootstrap policies from reference solutions. Also builds and evaluates reference-injection techniques for tasks such as code optimization and analyzes their effects on sample efficiency, bias, and robustness of policy updates.

reference-guidedpolicyimprovement

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing LLM instruction-tuning algorithms—such as supervised fine-tuning (SFT), proximal policy optimization (PPO), and direct preference optimization (DPO)—are often explained with heavy reliance on prior knowledge, omit critical derivations, or remain overly abstract, resulting in high cognitive barriers and poor interpretability. Method: This paper systematically unifies mainstream reinforcement learning and preference optimization approaches under a concise, symbolically grounded derivation framework explicitly tailored to practical LLM training scenarios. Contribution/Results: We introduce GRAPE (Generalized Relative Advantage Policy Evolution), a novel paradigm for future preference learning designed to overcome fundamental limitations of current methods in objective design, training stability, and generalization. The framework provides a coherent, step-by-step exposition—from SFT through DPO—enhancing algorithmic intuition and theoretical transparency. It establishes a rigorous foundation for advancing preference-based LLM alignment and offers principled directions for subsequent research.

Explaining reinforcement learning algorithms for instruction tuningIntroducing new research directions with GRAPE frameworkProviding clear intuitive understanding of complex RL methods

This work addresses the challenge of policy optimization in non-verifiable tasks, where explicit correctness signals are absent and existing group-wise comparison methods struggle to apply. The authors propose Reference Relative Policy Optimization (RRPO), a framework that constructs positive and negative anchor sets through hierarchical conditional trajectories and employs a set-based contrastive learning objective to train a metric projection head. This approach generates contrastive advantage scores without requiring a ground-truth verifier and integrates a group-wise centered policy update mechanism. RRPO effectively extends relative policy optimization to non-verifiable settings, achieving performance on par with or superior to verifier-based methods across verifiable reasoning, open-ended generation, and post-supervised fine-tuning scenarios, while significantly outperforming weakly supervised baselines.

contrastive comparisonnon-verifiable feedbackreinforcement learning

This work addresses the vulnerability of diffusion-based behavioral cloning policies to covariate shift, where minor state deviations can lead to catastrophic task failure. To mitigate this, the authors propose ReGuide, a novel framework that, for the first time, leverages guidance-generated successful trajectories at test time as reusable online recovery data to iteratively refine the policy during training. The core innovation is a Phase-Conditioned Guidance (PCG) mechanism that precisely synthesizes corrective trajectories within recoverable state regions. ReGuide enhances the base policy either through fine-tuning (ReGuide-FT) or full retraining (ReGuide-FS). Evaluated on multiple Robomimic tasks, ReGuide improves success rates by 1.3–7.7× over the baseline and substantially outperforms purely test-time guidance methods like LPB. Ablation studies confirm that the performance gains stem directly from the guided recovery data itself.

behavior cloningcovariate shiftdiffusion policies

Existing reinforcement learning approaches struggle to disentangle initial code generation quality from iterative self-repair capabilities in multi-turn code generation and often overlook intermediate execution signals. This work proposes TaPR, a framework that introduces a unified multi-turn interaction protocol to transform execution feedback into fine-grained test-passing-rate rewards, enabling the first decoupled evaluation of initial generation and self-repair performance. TaPR incorporates a reward decomposition mechanism and a turn-aware evaluation protocol, optimized through a dense reward strategy based on test pass rates. Experiments demonstrate that TaPR improves the three-turn pass rate (Pass@3) by 2.44 percentage points on LiveCodeBench and boosts accuracy from 30.25% to 33.56% on the 7B/8B model subset, significantly outperforming baseline methods.

code generationexecution feedbackmulti-turn interaction

This work addresses a critical limitation in existing reinforcement learning post-training methods—the absence of mechanisms to validate the efficacy of policy updates, which often leads to optimization drift or collapse. To mitigate this, the authors propose the PIRL framework, which reframes the optimization objective from immediate reward maximization to cumulative policy improvement across training iterations. Central to this framework is the PIPO algorithm, which introduces, for the first time, a policy improvement feedback mechanism. By employing a sliding window to retrospectively validate historical baselines, PIPO establishes a self-correcting closed-loop optimization process that guarantees each policy update positively contributes to final performance. Empirical evaluations on mathematical reasoning benchmarks demonstrate that PIPO achieves superior stability and performance compared to GRPO and its variants.

Large Language ModelsOpen-loop OptimizationPolicy Improvement

Latest Papers

What's happening recently
View more

This work addresses the challenge of sparse trajectory information in long-horizon reinforcement learning for large language model (LLM) agents, where weak policies often fail repeatedly, hindering effective policy optimization. The authors propose a policy-centric training paradigm that dynamically models skills as evolving scaffolds aligned with policy development. Specifically, during inference, the framework adaptively provides guidance through evidence card generation, task-specific evaluation, and context-aware adjustment mechanisms, gradually reducing reliance on external support as the agent’s capabilities improve—thus balancing guided assistance with growing autonomy. Integrated with standard RLVR optimization, this approach outperforms strong baselines by up to 18.6% on ALFWorld and WebShop benchmarks, achieves competitive performance across seven retrieval-augmented question-answering tasks, and reduces prompt usage by 32.1%.

LLM agentslong-horizon reinforcement learningpolicy optimization

Existing post-training methods in reinforcement learning lack dynamic, fine-grained control over the timing and location of interventions during rollout generation, often resulting in redundant or inefficient learning signals. This work proposes RAIL, a novel framework that introduces, for the first time, a recoverability-aware mechanism to model intervention decisions as an online contextual bandit problem. RAIL employs a shadow-to-live pipeline to collect intervention trajectories and leverages a critic-free population-based reinforcement learning architecture to train a recoverability controller that continuously adapts alongside the evolving policy. This controller autonomously determines when and where to intervene, optimizing rollout quality under limited computational budgets. Empirical results demonstrate that RAIL significantly enhances post-training efficiency across diverse scenarios, yielding high-quality rollouts with greater informational content and reduced redundancy.

adaptive decision-makingpost-trainingrecoverability-aware learning

This work addresses the vulnerability of large language model (LLM) agents to indirect prompt injection attacks, a threat inadequately mitigated by existing defenses due to their static or non-adaptive nature. The paper introduces the first structured feedback–driven iterative attack framework, which employs a rule-based diagnostic module to generate behavioral labels, an LLM-based optimizer that refines attack payloads using full interaction histories, and a seed synthesis mechanism that evolves attack strategies by learning from failed attempts. Integrating causal intervention and attention analysis, the study uncovers—for the first time—a threshold phenomenon in attention weights correlated with attack success. Experimental results demonstrate that the proposed framework substantially outperforms both static and state-of-the-art adaptive attacks on AgentDojo and InjectAgent benchmarks, achieving complete compromise on 5 out of 9 targets within the heavily defended Claude Code agent and significantly improving attack efficacy on the remaining targets.

Adversarial AttacksExternal Content VulnerabilityIndirect Prompt Injection

This work addresses the challenge that large language model agents struggle to dynamically refine their runtime execution frameworks using failure trajectories during deployment. To overcome this limitation, the authors propose training a dedicated “harness engineer” model that performs end-to-end, failure-conditioned editing of executable harnesses throughout their lifecycle. This is achieved through an initial cold-start supervised fine-tuning phase followed by online reinforcement learning via group relative policy optimization. The approach uniquely formulates harness editing as a learnable behavior and enables co-evolution with the target agent. Experimental results demonstrate that, on WebShop, ALFWorld, and DBBench, the method improves the base task success rate of Qwen3.5-9B from 44.3% to 53.6%, and further to 64.2% when combined with agent fine-tuning.

agent failure trajectoriesexecutable runtime harnessharness editing

Hot Scholars

ZF

Zoe Falomir

Umeå University, Sweden
Spatial CognitionQualitative ReasoningCognitive SystemsKnowledge-based
ZF

Zhenxing Fan

University of Virginia
Computer Architecture
YG

Yimin Gao

University of Virginia
AI hardwareProcessing-in-Memoryhardware/software codesignVLSI
LL

Ling Liu

Georgia Institute of Technology
Distributed Computing SystemsDatabase SystemsPrivacy-Security-TrustCloud Computing
MK

Mariam Kiran

Oak Ridge National Laboratory
networkingdistributed computingswarm agent-based modelsAI/ML