Mathematical Transfer in LLMs Follows Reasoning Approach More Than Topic

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether aligning training data by reasoning method or by mathematical topic better facilitates knowledge transfer during the fine-tuning of large language models for mathematics. Employing a balanced experimental design, we conduct 40 comparative fine-tuning experiments across multiple base models, validated through embedding similarity analysis and statistical significance testing. Our empirical findings reveal that cross-topic data sharing reasoning methods consistently outperforms same-topic data employing different methods, yielding average improvements exceeding 10 percentage points with confidence intervals strictly excluding zero. Based on these results, we establish reasoning method as the primary criterion for data curation. This finding challenges the conventional paradigm of organizing training data by domain classification, offering a new framework for constructing mathematical instruction datasets.
📝 Abstract
When selecting mathematical training data for LLMs, a natural organizing principle is topic: probability examples for probability targets. An alternative is reasoning approach: worked solutions that share a solution method with the target, even when the mathematical domain differs. We ask which relation produces greater transfer after fine-tuning. We evaluate two counterbalanced $2\times2$ designs: probability and combinatorics crossed with invariant reasoning and double counting (2,000 problems), and number theory and geometry crossed with complement and pigeonhole reasoning (800 problems). In each design, every cell serves as the held-out target in turn: same-approach (SA) sources share the target's method but change the topic, while same-topic (ST) sources share the topic but change the method. Every source appears once in each role, so additive source-quality effects cancel from the equally weighted aggregate contrast. Across five base models and three training seeds per design, SA outperforms ST in all 40 seed-pooled model--target comparisons. Model-level advantages range from 8.2 to 16.2 percentage points in the primary design (mean: 10.8) and from 12.0 to 16.0 in the second design (mean: 14.3); all ten model-level 95% confidence intervals exclude zero. In both designs, ST sources are more similar to targets under embedding and lexical measures, so the SA advantage runs opposite to the measured ordering of statement-level resemblance. These findings identify reasoning approach as a more effective matching criterion than topic for mathematical transfer across the evaluated topic--approach combinations.
Problem

Research questions and friction points this paper is trying to address.

Mathematical Transfer
Large Language Models
Reasoning Approach
Fine-tuning
Topic Matching
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mathematical Transfer
Reasoning Approach
Fine-tuning
Large Language Models
Cross-domain Generalization
🔎 Similar Papers
No similar papers found.