The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reasoning bottlenecks large language models encounter due to their lack of structured mathematical understanding. To this end, it pioneers the conceptual framework of "mathematical primitives" and constructs a four-dimensional diagnostic benchmark to precisely identify model deficiencies. Building upon this foundation, we propose a primitive-privileged self-distillation framework that leverages probing techniques and targeted repair mechanisms to achieve fine-grained, primitive-based knowledge distillation. Experimental results demonstrate that our approach significantly enhances the mathematical reasoning capabilities of models across varying scales, effectively overcoming identified bottlenecks and comprehensively outperforming existing baselines. This work establishes a novel paradigm for augmenting the structured mathematical reasoning of large language models.
📝 Abstract
While Large Language Models (LLMs) have demonstrated striking capabilities on frontier mathematical problems, it remains unclear whether they possess the structural mathematical understanding underlying their solutions. In this paper, we take a first step toward systematically studying mathematical understanding in LLMs, from diagnosing its distinct capabilities to leveraging these findings to improve post-training. First, we introduce the notion of Mathematical Primitive to probe structural mathematical understanding and propose \hlei{}, a novel benchmark that evaluates mathematical reasoning along four distinct dimensions: Discovery, Generation, Digestion, and Execution. Second, our systematic diagnosis shows that solution accuracy masks distinct capability profiles, primitives unlock substantial latent execution capacity, and Discovery is the dominant bottleneck in mathematical reasoning. Our post-training analysis further shows that discovery-limited failures are particularly amenable to repair. Finally, building on these findings, we introduce \abs{}, a primitive-privileged self-distillation framework that selectively transfers primitive-guided reasoning into the student model. Extensive experiments demonstrate that \abs{} consistently improves mathematical reasoning over baselines across model scales and challenging benchmarks.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Mathematical Reasoning
Structural Mathematical Understanding
Mathematical Primitive
Capability Diagnosis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mathematical Primitive
Self-Distillation
Mathematical Reasoning
Benchmark Evaluation
Large Language Models
🔎 Similar Papers
No similar papers found.