Solving Every Step Is Not Enough: Milestone Oracles Reveal a Composition Gap in LLM Math Reasoning

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the compositional gap in large language models during multi-step mathematical reasoning, where overall failure frequently occurs despite mastery of individual intermediate steps. Identifying this gap as a core bottleneck for the first time, we propose OracleLadder, a diagnostic framework that establishes a hierarchical fault attribution system to precisely localize reasoning breakpoints. The framework leverages teacher-generated fixed roadmaps, deterministic symbolic verifiers, and multi-level oracle assistance. Experiments across 354 problems demonstrate that the compositional gap accounts for 33–48% of failures, revealing significant discrepancies between step-wise accuracy and error recovery capabilities. Furthermore, our approach generalizes effectively to code generation, with findings successfully replicated on the MATH500 benchmark. All associated data and code have been made publicly available.
📝 Abstract
Large language models (LLMs) can solve every intermediate step of a multi-step math problem on its own and still fail the full problem, even when given a roadmap of the steps and all of their answers. We introduce OracleLadder, a diagnostic evaluation that locates where LLM math reasoning fails by giving the model increasing levels of oracle help. For each problem, a teacher model writes a fixed roadmap of intermediate sub-goals (milestones), and a deterministic symbolic verifier grades every answer. Testing the model with no help, with the roadmap, with the roadmap plus the milestone answers, and on each milestone alone sorts each failure into one of five reasoning gaps. On 354 NuminaMath problems and six models from 8B to 671B parameters (Qwen3, gpt-oss, Llama 3.3, DeepSeek-V3.1), the largest gap for every model is the composition gap, a stricter form of the compositionality gap. It covers 33-48% of problems, and 24-37% after removing problems that an LLM review flags as grading errors. Accuracy and milestone-help recovery rank the two strongest models differently, and two RLVR runs with similar accuracy gains move problems differently. The roadmap effect replicates on MATH500 and AIME 2024/25, per-problem recovery agrees for 83-87% of problems under an independent second teacher, and the help ladder carries over to code generation. We release the data, roadmaps, prompts, and code at https://github.com/slark-prime/OracleLadder.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Math Reasoning
Composition Gap
Multi-step Problem Solving
Milestone Oracles
Innovation

Methods, ideas, or system contributions that make the work stand out.

OracleLadder
Composition Gap
Milestone Oracles
Diagnostic Evaluation
Math Reasoning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Zhuohan Wang
Zhuohan Wang
Kings College London
Quantitative FinanceGenerative ModelGame Theory
Haoran Ma
Haoran Ma
PhD Student, University of California, Los Angeles
Computer SystemsSoftware Engineering
T
Tianyu Wu
Harvard University
Y
Yuanlin Duan
Rutgers University
Z
Zichun Liao
Harvard University
J
Jieming Yu
The Hong Kong University of Science and Technology