🤖 AI Summary
This study addresses the compositional gap in large language models during multi-step mathematical reasoning, where overall failure frequently occurs despite mastery of individual intermediate steps. Identifying this gap as a core bottleneck for the first time, we propose OracleLadder, a diagnostic framework that establishes a hierarchical fault attribution system to precisely localize reasoning breakpoints. The framework leverages teacher-generated fixed roadmaps, deterministic symbolic verifiers, and multi-level oracle assistance. Experiments across 354 problems demonstrate that the compositional gap accounts for 33–48% of failures, revealing significant discrepancies between step-wise accuracy and error recovery capabilities. Furthermore, our approach generalizes effectively to code generation, with findings successfully replicated on the MATH500 benchmark. All associated data and code have been made publicly available.
📝 Abstract
Large language models (LLMs) can solve every intermediate step of a multi-step math problem on its own and still fail the full problem, even when given a roadmap of the steps and all of their answers. We introduce OracleLadder, a diagnostic evaluation that locates where LLM math reasoning fails by giving the model increasing levels of oracle help. For each problem, a teacher model writes a fixed roadmap of intermediate sub-goals (milestones), and a deterministic symbolic verifier grades every answer. Testing the model with no help, with the roadmap, with the roadmap plus the milestone answers, and on each milestone alone sorts each failure into one of five reasoning gaps. On 354 NuminaMath problems and six models from 8B to 671B parameters (Qwen3, gpt-oss, Llama 3.3, DeepSeek-V3.1), the largest gap for every model is the composition gap, a stricter form of the compositionality gap. It covers 33-48% of problems, and 24-37% after removing problems that an LLM review flags as grading errors. Accuracy and milestone-help recovery rank the two strongest models differently, and two RLVR runs with similar accuracy gains move problems differently. The roadmap effect replicates on MATH500 and AIME 2024/25, per-problem recovery agrees for 83-87% of problems under an independent second teacher, and the help ladder carries over to code generation. We release the data, roadmaps, prompts, and code at https://github.com/slark-prime/OracleLadder.