AutoResearch at Production Scale: Failure Modes and a Multi-Agent Framework

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive costs of embedding optimization and the structural failures inherent in automated exploration within production-grade recommendation systems by extending the AutoResearch paradigm to industrial scale. Through identifying five application-agnostic failure modes, this work proposes a multi-agent framework adhering to “prevent, persist, and redirect” principles. By integrating LLM-driven code editing, multi-GPU clusters, and automated evaluation mechanisms, it constructs an iteration-cost-scaling scaffolding design to stabilize the research pipeline. Experimental results demonstrate that the proposed approach achieves a 1.82× improvement in Recall@6 and a 2.1× increase in consistency, while an autonomous text fallback mechanism expands catalog coverage by 5.8×.
📝 Abstract
Optimizing embedding systems for production recommendation pipelines demands systematic exploration that consumes disproportionate engineering effort at scale. We apply Andrej Karpathy's AutoResearch paradigm -- a large language model that iteratively edits a training script and retains modifications that improve a held-out scalar metric -- to automate this exploration. We report on twelve weeks of running this paradigm at production scale, where iterations consume hours of multi-GPU compute, evaluation involves competing criteria, and campaigns span weeks across many training jobs. Across two independently developed representation-learning systems for a book recommendation pipeline, we ran 220+ experiments and observed five recurring failure modes absent from the original setting: infrastructure fragility, agent memory decay, search-direction stagnation, iteration-cost asymmetry, and metric fixation. We contribute a three-principle scaffolding design -- prevent, persist, redirect -- that maps each failure mode to a structural remedy and whose instantiation scales with iteration cost. The framework produced a 1.82x Recall@6 lift and a 2.1x coherence lift over hand-tuned baselines, and the agent autonomously designed a text-only fallback that expanded catalog coverage by 5.8x. The two systems span nearly three orders of magnitude in per-iteration cost yet exhibit the same failure modes, suggesting these are structural properties of production-scale autonomous research rather than artifacts of either application.
Problem

Research questions and friction points this paper is trying to address.

AutoResearch
production-scale recommendation
embedding optimization
failure modes
multi-agent framework
Innovation

Methods, ideas, or system contributions that make the work stand out.

AutoResearch
Multi-Agent Framework
Production-Scale Recommendation
Failure Modes
Representation Learning