🤖 AI Summary
This study addresses the limitation of scalar feedback in automated large language model (LLM) research, where it fails to reveal behavioral conflicts that cause performance stagnation. To overcome this, we propose a competitive behavior feedback mechanism that identifies such conflicts and designs probe metrics to guide code agents in optimizing models. Furthermore, we introduce a reusable ConflictGuide-Skill that integrates literature taxonomy with model evidence to conduct a two-stage evolutionary search for conflict mitigation. Experimental results demonstrate that our approach reduces task error rates by 28% and conflict-related error rates by 14% across five model families, significantly outperforming baseline methods that rely solely on scalar feedback.
📝 Abstract
When designing machine learning models, desirable properties are often in tension: improving one behavior can impair another, so task progress can depend on alleviating the conflict. LLM-based AutoResearch systems, which iteratively edit model code and retain edits based on scalar task-performance feedback, have largely ignored this trade-off. We find that scalar feedback supports broad exploration early in search, but it does not reveal how edits affect competing behaviors. In matched-budget experiments, introducing competing-behavior feedback as task gains diminish increases the share of proposals that improve both behaviors and sustains progress beyond scalar-only plateaus. Obtaining this feedback for a given model requires identifying its competing behaviors and designing probes to measure them. To make competing-behavior feedback actionable, we introduce ConflictGuide. Its reusable ConflictGuide-Skill combines a literature-grounded taxonomy with model-specific evidence to identify competing behaviors and specify probes for a code agent to implement as metrics. Evolution proceeds in two stages: Stage I explores with task feedback; Stage II uses probe feedback to steer proposals toward conflict alleviation and retains marginal-gain edits only when probes indicate sufficient alleviation. Across five diverse model families, ConflictGuide reduces task and conflict-related errors by up to 28% and 14%, respectively, relative to scalar-only AutoResearch, with gains extending to other code agents.