RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the problem of silent failures, such as execution deviations and data leakage, that compromise experiments in long-horizon automated research. To this end, we propose a multi-agent automatic evolution framework that integrates state-machine-based runtime control, executable action protocols, and the HSTU recommendation architecture. Furthermore, we introduce a novel meta-meta-testing mechanism in which heterogeneous black-box code agents collaboratively review and repair artifacts to ensure reliable knowledge transfer across iterations. Experimental results demonstrate that the proposed system improves execution accuracy to 62.5% and increases NDCG@10 by 4.48% on the MovieLens benchmark. These findings effectively validate the reliability and superiority of heterogeneous agent collaboration for automated scientific research.
📝 Abstract
Auto-research agents, LLM systems that propose, implement, train, and evaluate model changes across iterations, promise to automate applied ML's experimental loop. Over long horizons, execution accuracy is a binding constraint: a change can silently leak held-out data, omit normalization, disconnect a gradient, or leave a train/eval flag unwired, invalidating expensive runs and compounding error across iterations. We present RankEvolve, an auto-research framework for evolving generative ranking models. An Executable Operating Protocol (EOP) declares phases, gates, branches, and loops, and the runtime enforces the compiled state machine. A meta-meta-harness composes complete black-box coding-agent products, including Claude Code and Codex, as execution-graph nodes that review and repair one another's work. In a budget-matched evaluation, heterogeneous composition raises all-oracle execution accuracy from the best single-product baseline of 45.8 percent to 62.5 percent (paired +16.7 points, 95 percent CI [6.6, 26.7]) while achieving a 10.4 percent silent critical-defect rate. An implemented knowledge layer carries findings, including negative results, across iterations. In a twelve-iteration deployment on the open-source HSTU recommender, RankEvolve reported NDCG@10 of 0.2192 on MovieLens-20M LARGE (+4.48 percent over the published anchor) and 0.1948 on BASE (+2.80 percent). ExecML-HSTU, seeded by incidents from that deployment, provides the oracle benchmark for the execution-accuracy evaluation. A pre-specified LitGPT transfer split replicates the heterogeneous-composition effect beyond recommendation (+12.5 points, 95 percent CI [3.0, 22.0]), and a paired ablation isolates per-step from full-protocol instruction injection. These results characterize when runtime-controlled composition of coding-agent products improves execution accuracy.
Problem

Research questions and friction points this paper is trying to address.

auto-research agents
execution accuracy
silent defects
iterative model evolution
automated machine learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Auto-research agents
Executable Operating Protocol
Multi-agent composition
Execution accuracy
Ranking models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zheng Chen
Linfeng Liu
Linfeng Liu
Meta
H
Hong Li
H
Hong Yan