The Imitation Game: When LLMs Learn to Reason Like Programs via Code-Centric Reasoning Data Synthesis

📅 2026-09-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决大语言模型在自然语言精细推理上的不足,提出MIMIC框架,通过代码合成数据进行训练,提升模型的广泛推理能力。
📝 Abstract
Large Language Models (LLMs) excel at programming tasks but frequently fail at deterministic, fine-grained reasoning in natural language, relying heavily on semantic approximations rather than robust symbolic execution. To bridge this gap, we propose MIMIC, a framework that leverages executable code as a rigorous medium for reasoning data synthesis. MIMIC fundamentally transforms algorithms into verifiable reasoning trajectories through narrative fusion, code-guided test synthesis, and dynamic code instrumentation. Crucially, these explicit intermediate execution states naturally form a Code-Instrumented Reward (CIR), providing dense, high-fidelity process supervision for reinforcement learning without external reward models. Extensive evaluations reveal that models trained via SFT and GRPO on our synthesized dataset achieve substantial, consistent gains. Our method significantly elevates accuracy across general reasoning, complex mathematical benchmarks, and fine-grained deterministic tasks, demonstrating that the procedural rigor of executable code can effectively unlock and enhance the generalized reasoning capabilities of LLMs. Our code and data are available at https://github.com/zjy1298/MIMIC.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
deterministic reasoning
fine-grained reasoning
symbolic execution
natural language
Innovation

Methods, ideas, or system contributions that make the work stand out.

MIMIC
Code-Instrumented Reward (CIR)
code-centric reasoning
reinforcement learning
executable code
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jinyang Zhang
School of Computer Science, Peking University
Weibin Liao
Weibin Liao
Peking University
Large Language ModelReinforcement LearningMedical Image Analysis
Keqin Bao
Keqin Bao
University of Science and Technology of China
Large Language ModelsRecommender Systems
S
Sihang Li
Qwen Team, Alibaba Group
Shaobo Wang
Shaobo Wang
Shanghai Jiao Tong University
Large Language ModelData-Centric AIData SynthesisData SelectionExplainable AI
M
Muyang Ye
College of Computer Science and Technology, Zhejiang University
H
Hongxin Ding
School of Computer Science, Peking University
Y
Yue Fang
School of Computer Science, Peking University
Tianyi Tang
Tianyi Tang
Qwen Team, Alibaba Group & Renmin University of China
Artificial IntelligenceNatural Language Processing
Fei Huang
Fei Huang
Qwen Team, Alibaba Group
Natural Language Generation
Kexin Yang
Kexin Yang
Qwen Team
Natural Language ProcessingControllable Text Generation
X
Xingzhang Ren
Qwen Team, Alibaba Group
D
Dayiheng Liu
Qwen Team, Alibaba Group