Feedback Over Form: Why Execution Feedback Matters More Than Pipeline Topology in 1-3B Code Generation

📅 2026-04-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates whether model ensembles within the 1–3B parameter range can enhance code generation performance through execution feedback and pipeline architectures. We construct a generate-and-refine pipeline based on small language models, incorporate an execution feedback mechanism, and employ a NEAT-inspired evolutionary algorithm to search for optimal topologies. Our experiments reveal that execution feedback is pivotal—yielding performance gains exceeding four standard deviations on HumanEval and MBPP, primarily by correcting runtime errors—whereas increased topological complexity offers no significant benefit. The refinement component’s capability outweighs the identity of the generator, and single-run evaluations tend to overestimate evolutionary improvements; early stopping proves essential to prevent performance degradation. Moreover, specialized code models consistently outperform all combinations of general-purpose models.

Technology Category

Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageSearch and Optimization: Evolutionary ComputationMachine Learning: Evolutionary Learning

Application Category

Economics, Online Markets and Human Computation: Economic ramifications for generative AI infrastructure and applicationsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsSocial Networks and Social Media: Generative AI / large language models and their impact on social systems
📝 Abstract
Small language models (1-3B) are practical to run locally, but individually limited on harder code generation tasks. We ask whether composing them into pipelines can recover some of that lost capability. We study code generation pipelines built from 1-3B models with execution feedback, and use a NEAT-inspired evolutionary search to test whether more complex pipeline structure helps beyond a simple refinement loop. We evaluate on HumanEval (164 problems) and sanitized MBPP (427 problems), all with local inference on a single laptop. Self-refinement with execution feedback improves code generation by more than 4 standard deviations on both benchmarks. The gains are narrow in mechanism: refinement fixes many runtime errors (especially NameError and SyntaxError), but rarely fixes logic errors such as AssertionError. Within our tested general-purpose model pool, generator identity mattered less than refiner capability: a 1.5B generator paired with a 3B refiner matched a 3B model doing both roles. Early stopping is essential; without it, every iteration is net-negative. The code-specialized models outperform every general-purpose pipeline configuration, suggesting model specialization matters more than pipeline architecture. Preliminary text-only pipeline experiments without execution feedback did not show gains at this scale. In our constrained search space, evolutionary search mostly rediscovered the same simple generate-execute-refine loop we found manually, with no clearly significant gain from added topology. Single-evaluation fitness inflates results by 5-7 percent, selecting lucky genomes over good ones. On these benchmarks at 1-3B scale, execution feedback mattered more than added pipeline complexity in determining whether composition helped.
Problem

Research questions and friction points this paper is trying to address.

code generation
small language models
execution feedback
pipeline topology
model composition
Innovation

Methods, ideas, or system contributions that make the work stand out.

execution feedback
code generation
small language models
pipeline topology
evolutionary search
C
Charles Junichi McAndrews