multi-step reasoning

Designs, implements, or evaluates algorithms, models, and systems that construct, execute, or verify chains of inferential steps to reach conclusions or plans, including decomposition into subproblems, management of intermediate representations, and control of step sequencing. This work also involves analyzing error propagation, stepwise correctness, and stopping or convergence criteria for multi-step procedures.

multi-stepreasoning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.28
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$199K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge that small language models often fail to reliably execute multi-step, dependency-rich structured graph algorithms due to error accumulation. The authors frame algorithm execution as a closed-loop prediction task, wherein the model iteratively selects operations based on the current graph state and evaluates its overall behavior through full rollbacks. Departing from conventional step-isolated evaluation, this closed-loop rollback paradigm reveals that strong single-step prediction accuracy does not necessarily ensure stable global execution. Experimental results demonstrate that suitably adapted small models can reliably perform algorithms such as traversal and coloring, yet remain vulnerable to cumulative errors in weighted graph algorithms. These findings underscore the necessity and efficacy of the proposed closed-loop evaluation framework for assessing and improving algorithmic reasoning in language models.

closed-loop executiongraph algorithmsrollout reliability

Error estimation and step size control with minimal subsystem interfaces

Jun 25, 2024
LT
Lars T. Kyllingstad
🏛️ SINTEF Ocean | SINTEF Ålesund

Industrial “black-box” subsystems—characterized by coupled outputs only, no rollback capability, absence of derivative information, and inaccessibility of internal states—pose fundamental challenges for error quantification and step-size control in co-simulation. Method: This paper proposes the first co-simulation error estimation and macro-step adaptation framework tailored to minimal interface constraints. It introduces a model-free error indicator based on local truncation error modeling and multi-step extrapolation comparison, designs a robust, configurable step-size adjustment algorithm, and provides pseudocode-level reproducible control logic. Contribution/Results: It is the first systematic solution to error quantification under purely output-coupled, zero-internal-information conditions. The framework significantly enhances accuracy controllability, computational efficiency, and engineering reliability in large-scale heterogeneous co-simulations. The work includes a comprehensive implementation guide and practical anti-pattern recommendations.

Controlling macro step sizes for speed-accuracy balanceEnabling black-box subsystem coupling in industrial applicationsEstimating co-simulation errors with minimal subsystem interfaces

This work addresses the challenge of cascading failures in multi-step reasoning, where a single erroneous step can compromise the entire solution. Existing model routing approaches treat all reasoning steps uniformly, leading to suboptimal efficiency. To overcome this, we propose TRIM, the first method enabling dynamic, step-level routing: leveraging a process reward model to identify high-uncertainty, critical steps, TRIM selectively delegates only these error-prone steps to a large language model under a computational budget, while simpler steps are handled by a smaller, more efficient model. This strategy effectively interrupts error propagation. Experiments demonstrate that TRIM achieves up to 5× higher cost efficiency than prior methods on MATH-500, with advanced routing strategies matching baseline performance using 80% fewer large-model tokens. On challenging benchmarks like AIME, TRIM delivers up to 6× cost efficiency gains while maintaining strong generalization.

cascading failuresLLM routingmathematical problem solving

This work addresses the lack of machine-verifiable formalizations of line search methods in nonlinear optimization, which has hindered algorithmic reliability. Within the Lean 4 theorem prover, it presents the first systematic formalization of several classical line search criteria—including Armijo, Goldstein, Wolfe, and their nonmonotone variants—alongside rigorous definitions of gradient descent, descent directions, and backtracking step-size selection. The study fully verifies the Zoutendijk convergence theorem within this framework, thereby establishing a comprehensive formal foundation for line search theory. This contribution significantly enhances the verifiability and trustworthiness of nonlinear optimization algorithms through mechanized mathematical reasoning.

convergenceformalizationline search

Outcome-Refining Process Supervision for Code Generation

Dec 19, 2024
ZY
Zhuohao Yu
🏛️ Peking University | Microsoft Research

Large language models (LLMs) exhibit insufficient reasoning capabilities for complex programming tasks: process supervision relies on costly and error-prone reward modeling, while outcome supervision struggles to coordinate multi-step reasoning. To address this, we propose a novel “outcome-refinement-as-process” supervision paradigm that eliminates explicit reward modeling and instead leverages program execution feedback—such as runtime outputs and error traces—as label-free, reliable intermediate supervision signals. Our approach integrates tree-based multi-path exploration with a lightweight model adaptation framework to enable efficient, execution-guided reasoning. Evaluated across five LLMs and three benchmark datasets, our method achieves average improvements of 26.9% in code correctness and 42.2% in execution efficiency. Notably, it significantly boosts the performance of smaller models on algorithmic competition–style tasks. This work establishes a scalable, low-overhead paradigm for complex programming reasoning, grounded in direct execution feedback rather than surrogate reward signals.

Improving code generation for complex programming tasksOvercoming local optima in LLM-generated codeUnifying process and outcome supervision via execution

Latest Papers

What's happening recently
View more

Although large language models excel at reasoning tasks, they often fail to faithfully execute multi-step procedures specified in prompts. This work introduces a controlled diagnostic benchmark comprising simple arithmetic programs with variable lengths and backtracking dependencies to systematically evaluate execution fidelity across 14 prominent models on 55 datasets. The study reveals characteristic failure modes in long programs—such as step omission, premature answering, erroneous self-correction, under-execution, and hallucinated steps—for the first time. While models achieve a 61% first-answer accuracy on 5-step programs, performance sharply declines to 20% on 95-step programs, underscoring their significant limitations in following complex, extended instructions.

diagnostic benchmarkinstruction followinglarge language models

This work proposes a formalization of algorithms within an intensional computability framework and clarifies their relationship to implementations in computational models. Treating computational models as monoid actions on configuration spaces, programs are modeled as dynamical systems constrained by such actions. Algorithms are defined as finite directed graphs of partial maps over edge-labeled abstract data structures, explicitly separating control flow from data operations. By leveraging tools from category theory, dynamical systems theory, and graph theory, the approach constructs a rigorous semantic framework that, for the first time, treats algorithms as abstract specifications of computational behavior and precisely characterizes the structure-preserving implementation relation between programs and algorithms, thereby deepening our understanding of the nature of computation.

abstract data structurealgorithmcomputability

This work addresses the undecidability of formal verification for model transformations, which stems from Turing completeness, and the path explosion problem that persists even in non-Turing-complete domain-specific languages like DSLTrans. The authors propose a scalable verification approach by establishing, for the first time, a bounded completeness theorem for a fragment of DSLTrans with respect to existential and traceability properties, thereby reducing infinite verification problems to bounded yet complete checks. Their method integrates class-boundary-aware encoding, trace-aware dependency analysis, and a CEGAR-driven refinement strategy to drastically reduce SMT formula size and eliminate spurious counterexamples. Implemented atop Z3 and integrated into a Web IDE, the tool successfully verifies 552 out of 899 properties across 29 real-world transformations, generates 345 valid counterexamples, times out on only two cases, and achieves up to a 112× speedup on challenging instances through refinement.

bounded model checkingDSLTransformal verification

Existing benchmarks struggle to evaluate the step-level process quality and error propagation of tool-using agents in open, dynamic environments. This work proposes the first step-level evaluation benchmark tailored to real-world tool invocation scenarios, comprising 1,000 execution trajectories and 8,509 human annotations with high inter-annotator agreement (IAA: 89.1%). The benchmark introduces ternary action labels—correct, neutral, and incorrect—and formalizes error propagation rules to construct a trajectory-level process evaluation framework. The study reveals that weak policy models exhibit inflated task success rates due to premature termination; current models have difficulty distinguishing neutral from erroneous actions; and incorporating process-level signals significantly enhances test-time generalization, thereby demonstrating their complementary value to outcome-based supervision.

error propagationopen-ended tool executionprocess quality

Hot Scholars

HZ

Hamed Zamani

Associate Professor of Computer Science, University of Massachusetts Amherst
Information RetrievalRecommender SystemsNatural Language ProcessingConversational AI
XJ

Xiaozhong Ji

Nanjing University
computer visionimage processingsuper resolution
ZD

Zhicheng Dou

Renmin University of China
Information RetrievalRetrieval Augmented GenerationLarge Language ModelsGenerative IR
FW

Furu Wei

Distinguished Scientist, Microsoft Research
Natural Language ProcessingArtificial IntelligenceGeneral AIGenerative AI
LY

Lingyong Yan

Baidu Inc.
Large Language ModelMachine Learning