MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the persistent challenge large language models face in synthesizing complete programs from scratch, primarily due to the absence of scalable training environments that span the full software engineering lifecycle. To overcome this limitation, the authors introduce MindForge, an automated pipeline that constructs the first source-code-free training environment by stripping source code from open-source command-line programs and retaining only executables and documentation. Leveraging a teacher model, they generate high-quality synthetic execution trajectories to supervise fine-tuning of a smaller model (Qwen3.6-27B). Evaluated on ProgramBench, this approach improves average test pass rates from 37.98% to 49.51% and demonstrates substantial gains across seven previously unseen software engineering benchmarks, with performance increases up to 31.00 percentage points, effectively enabling smaller models to approach the program synthesis capabilities of state-of-the-art large models.
📝 Abstract
Coding agents have made substantial progress on software engineering tasks that modify existing codebases, including bug fixing and feature implementation. However, constructing a complete program from scratch remains a major challenge: even the frontier models evaluated on ProgramBench fully resolve fewer than 1% of tasks. One obstacle is the lack of scalable training environments for this from-scratch setting, spanning the whole software engineering life cycle, as existing environment-construction frameworks focus only on a single phase in software development. To address this gap, we introduce MindForge, an automated pipeline that converts open-source command-line programs into source-free environments that expose only a compiled reference executable and its documentation. Using MindForge, we construct training environments from repositories disjoint from those in ProgramBench, and curate a high-quality data recipe consisting of program synthesis trajectories using GLM-5.2 as the teacher agent. Fine-tuning Qwen3.6-27B on these trajectories increases its ProgramBench average test pass rate from 37.98% to 49.51%, achieving performance comparable to substantially larger frontier models. Moreover, the fine-tuned model consistently improves over the base model across all seven unseen software engineering benchmarks, spanning long-horizon repository generation and translation, bug fixing, feature implementation, and cross-language issue resolution, with absolute gains of 31.00 points on RepoZero-C2Rust, 14.16 on DeepSWE, 10.70/4.56 on NL2Repo-Bench (with/without tests), 5.04 on SWE-bench Verified, 5.93 on SWE-bench Pro, 5.22 on SWE-bench Multilingual, and 4.94 on FeatBench.
Problem

Research questions and friction points this paper is trying to address.

program synthesis
software engineering lifecycle
source-free environments
from-scratch programming
training environment
Innovation

Methods, ideas, or system contributions that make the work stand out.

source-free program synthesis
whole-life-cycle software engineering
MindForge
program synthesis trajectories
coding agent fine-tuning
🔎 Similar Papers