🤖 AI Summary
This study addresses the inefficiency of model iteration in industrial recommender systems caused by reliance on manual decision-making. To this end, we propose an autonomous research framework driven by the collaboration of dual Research and Model agents. This approach constructs a business sandbox environment that bridges proposals and experiments through a four-step action loop—including reproduction and follow-up—thereby automating the entire pipeline from problem diagnosis to code verification. Across 560 experiments, the proposed system outperforms baseline methods, achieving a 10–20% improvement in efficiency and reducing computational costs by up to 10% for certain models. Furthermore, it reveals that complex scheduling mechanisms are unnecessary during preliminary analysis stages, validating both the feasibility and effectiveness of long-horizon autonomous model research.
📝 Abstract
Sustaining industrial recommendation research requires using the results of one experiment to decide what to investigate next. We present AgentX-Model, the next generation of AgentX's model research framework, which connects proposal development and model experimentation within sandboxes defined by business inputs and prediction tasks. AgentX-Model adopts a dual-agent architecture comprising a Research Agent and a Model Agent. The Research Agent develops independently reviewed proposals from papers and experimental findings, while the Model Agent conducts multi-round investigations and returns code, measurements, and unresolved questions. Using the returned results, the Research Agent selects a starting implementation and formulates the next research question, allowing subsequent experiments to build on earlier findings. We organize this continuing research around four actions: Reproduce, Follow-up, Composition, and Diagnose. The first three actions drive routine research, while Diagnose acquires the evidence needed to choose a repair, including for issues raised by business feedback and online evaluation, such as prediction bias measured by PCOC. Across the production evaluation, 560 of 636 completed model-changing experiments recorded AUC above their business baselines. As research continued, some experiments recorded AUC above every comparable ancestor in their lineages. The five latest online A/B evaluations across different business settings reported gains including 10-15% in acquisition efficiency, 15-20% in target-segment advertising spend, and 0.3-0.8% in watch time; the watch-time model used approximately 10% fewer FLOPs and parameters. A dependency-aware historical-replay benchmark further evaluates research allocation, with initial results showing no consistent efficiency gain from more complex scheduling when agents already analyze and select concrete candidates.