🤖 AI Summary
This work addresses the challenge of continuously optimizing user-defined reasoning scaffolds at low cost within a framework of co-evolution between models and inference scaffolds, thereby enhancing the quality of agent execution trajectories. It proposes a recursive scaffold self-improvement mechanism that unifies scaffold optimization with model training data generation, representing scaffolds via prompt-level specifications and iteratively refining them through pairwise preference feedback derived from their own revision history. Rather than extending reasoning chains, the method emphasizes task context management to achieve efficient information flow control with minimal inference overhead. Evaluated on 30 cross-domain synthetic tasks, the approach significantly outperforms high-overhead baselines within just a few iterations, reducing inference costs by up to 60%, thus demonstrating the critical role of effective context management in performance improvement.
📝 Abstract
Under model--harness co-evolution, harnesses are not merely inference-time scaffolds but data-generating components whose execution traces can shape future foundation models. This motivates harness-in-the-loop learning: optimizing harnesses for both immediate agent performance and the quality of traces used for future model training. However, continually updating provider-built scaffolds is costly and labor-intensive. We therefore investigate whether optimizing user-constructed harnesses in a task-specific manner can improve execution-trace quality while remaining computationally lightweight and requiring only a few update iterations. To this end, we introduce Recursive Harness Self-Improvement (RHI), which represents the harness as a prompt-level specification of the agent loop and iteratively refines it using pairwise feedback over its own revision history. Across 30 synthetic machine-learning research tasks spanning quantitative finance, robotics, and pharmacy, a few RHI iterations suffice to substantially raise the performance ceiling of low-reasoning-effort agents, exceeding the corresponding maximum-reasoning-effort setting while reducing inference cost by up to 60%. We show that these gains arise primarily from improved task-specific context management through more effective inter-agent information flow rather than longer reasoning traces. Finally, we formalize this behavior as an information-theoretic hypothesis for RHI's implicit optimization objective, suggesting RHI as a practical algorithm for continual learning within the paradigm of model--harness co-evolution.