Score
Designs and implements closed-loop rollout simulations and multi-rollout evaluation pipelines that execute policies from prefixes through full trajectories to measure behavior under feedback and partial solutions. Builds execution diagnostics and analyses — e.g., exact rollout accuracy, prefix survival, constraint validity, and local decision quality — and uses intervention-based experiments to locate and quantify error accumulation during execution.
This work addresses the challenge that small language models often fail to reliably execute multi-step, dependency-rich structured graph algorithms due to error accumulation. The authors frame algorithm execution as a closed-loop prediction task, wherein the model iteratively selects operations based on the current graph state and evaluates its overall behavior through full rollbacks. Departing from conventional step-isolated evaluation, this closed-loop rollback paradigm reveals that strong single-step prediction accuracy does not necessarily ensure stable global execution. Experimental results demonstrate that suitably adapted small models can reliably perform algorithms such as traversal and coloring, yet remain vulnerable to cumulative errors in weighted graph algorithms. These findings underscore the necessity and efficacy of the proposed closed-loop evaluation framework for assessing and improving algorithmic reasoning in language models.
This work addresses the lack of closed-loop control in traditional software development lifecycles, which often fails to simultaneously ensure security, auditability, and highly reliable automation. The authors propose a deterministic autonomous control framework that models the lifecycle as a seven-stage automated pipeline, integrating Jira-based task orchestration, structured context, resource constraints, and human-review gating mechanisms to establish a secure closed loop. Key innovations include a state-contract-based collision locking mechanism, a degradation protocol for fallback operation, and a traceable control architecture. Implemented with 12,661 lines of Python code and 6,907 lines of versioned prompt specifications—including 101 exception handlers and 12 centralized locks—the system achieved a 100% success rate (95% CI [97.6%, 100%]) across 152 initial runs, producing over 795 artifacts. All 51 issues identified through adversarial review were fully resolved, with 60% of security tickets autonomously completed.
A long-standing debate in reinforcement learning for large language models concerns the relative statistical difficulty of process supervision versus outcome supervision, with conventional wisdom favoring the former. Method: This paper theoretically analyzes their statistical complexity under standard data coverage assumptions and introduces the trajectory measure transformation lemma—a novel tool that formally links return-oriented trajectory distributions to step-level distributional shifts. It further proves that any policy’s advantage function serves as an optimal process reward model. Results: The work establishes statistical equivalence between process and outcome supervision, demonstrating that observed performance gaps stem from algorithmic implementation flaws—not intrinsic statistical hardness. By unifying the theoretical foundations of both paradigms, it provides rigorous justification for lightweight, high-efficiency outcome supervision, challenging the prevailing reliance on fine-grained process annotations.
This work proposes a novel paradigm that integrates discrete-event simulation with a large language model (Gemini-1.5-Pro) to overcome the limitations of traditional simulation-based optimization, which treats simulators as black boxes and offers little insight into policy failure. By leveraging event-level trajectory replay, the method automatically identifies bottlenecks from low-scoring simulation runs and generates interpretable, traceable, code-level policy revisions in parallel. It pioneers the use of simulation trajectories to guide the LLM in targeted heuristic rule modification, combined with rolling evaluation and an elite retention mechanism for iterative policy improvement. Evaluated on dynamic production and AGV scheduling tasks, the approach achieves an average policy score of 77.51 (out of 100), improving the best run from 62.49 to 78.61, and significantly outperforms MILP, handcrafted rules, and metaheuristic baselines across 100 random seeds and fault perturbations.
This work addresses the high computational cost of autoregressive rollouts and the issues of policy lag and distribution shift arising from multi-epoch reuse in large language model reinforcement learning. To overcome these challenges, the paper proposes Prefix-Normalized Policy Optimization (PNPO), which innovatively replaces conventional cumulative importance weights with the geometric mean of likelihood ratios computed over causal prefixes. This approach preserves prefix dependencies while compressing the dynamic range of importance weights, enabling stable and efficient off-policy updates. Experimental results demonstrate that under a four-epoch update setting, PNPO achieves an Avg@32 of 50.24% across multiple mathematical reasoning benchmarks, outperforming GSPO by 3 percentage points. Moreover, within the same computational budget, PNPO attains performance equivalent to that of single-epoch training with 600 rollout batches using only 150 batches.
Existing reinforcement learning approaches struggle to disentangle initial code generation quality from iterative self-repair capabilities in multi-turn code generation and often overlook intermediate execution signals. This work proposes TaPR, a framework that introduces a unified multi-turn interaction protocol to transform execution feedback into fine-grained test-passing-rate rewards, enabling the first decoupled evaluation of initial generation and self-repair performance. TaPR incorporates a reward decomposition mechanism and a turn-aware evaluation protocol, optimized through a dense reward strategy based on test pass rates. Experiments demonstrate that TaPR improves the three-turn pass rate (Pass@3) by 2.44 percentage points on LiveCodeBench and boosts accuracy from 30.25% to 33.56% on the 7B/8B model subset, significantly outperforming baseline methods.