🤖 AI Summary
This study addresses the challenges of cross-task transfer and high continuous optimization costs in LLM-driven model evolution by proposing a population-based evolutionary framework. The method introduces a collaborative hierarchical evolution mechanism that connects local searches through shared empirical evidence, enabling efficient scaling from small-scale screening to large-scale training. By integrating code-level mutation, peer evaluation, and hierarchical population management, alongside the proposed RMD-Bench benchmark, the framework facilitates effective knowledge sharing across tasks. Experimental results demonstrate that this approach significantly outperforms independent evolution in ranking, watch time prediction, and large model pre-training, achieving accuracy improvements of up to 2.48%. Overall, the proposed framework realizes efficient computational resource allocation while promoting robust cross-task knowledge transfer.
📝 Abstract
LLM-driven evolution enables iterative model development, but two practical goals remain underexplored: finding model designs that transfer across related tasks and sustaining improvement when training is expensive. We introduce Population Evolution (PE), a collaborative, hierarchical framework that connects ongoing local searches through shared experimental evidence. PE evaluates code changes across related training instances and shares the results to guide subsequent proposals and promotion to larger training scales. For expensive targets, PE searches small training subsets and screens candidates through peer and intermediate evaluations before full-target training. We introduce RMD-Bench to evaluate both settings across ranking, watch-time prediction, RL algorithm discovery, and LLM/VLM pretraining. Compared with standalone evolution at matched source iterations, PE raises mean best local gains from 7.01% to 8.97% in ranking and from 2.84% to 3.85% in watch-time, while improving the best larger-scale outcome in all three joint-discovery families. In watch-time discovery, PE improves best larger-scale gains with four of five harnesses and all four proposers. On new recommendation datasets under shared target-side calibration, every evaluated PE design improves over the reference in mean performance. Under matched total GPU compute, completed LLM discovery runs yield a best relative accuracy gain of 2.48% and 13 successful candidates for PE, versus 0.92% and none for direct evolution. VLM loss reduction reaches 8.78% versus 5.05% under matched total GPU compute.