Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work presents the first practical implementation of recursive self-improvement (RSI) in machine learning engineering, introducing a unified learn-and-evolve loop framework. The authors develop OpenMLE, a full-stack executable AI4AI system that integrates a verifiable task environment, operator learning, and long-horizon search modules. They apply execution-driven supervised fine-tuning and reinforcement learning to the Frontis-MA1 (35B) model, transforming it into a meta-evolutionary agent that performs closed-loop optimization through four atomic operations: Draft, Improve, Debug, and Crossover. Leveraging deduplicated training data and an asynchronous experience prior mechanism (OpenMLE-Evo-Max), the system achieves a Medal Average of 71.21% on MLE-Bench Lite—surpassing GPT-5.5+Codex—and attains a 70% Match-SOTA score on NatureBench Lite, demonstrating strong generalization and cross-domain capabilities.
📝 Abstract
Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a concrete, executable testbed for studying this capability. We introduce OpenMLE, an open full-stack system for RSI research in MLE, spanning verifiable task environments with execution feedback (OpenMLE-Gym), operator learning (OpenMLE-RL), and long-horizon search (OpenMLE-Evo). On this stack we post-train Frontis-MA1 (35B) as a meta-evolution agent for MLE, aligning post-training and inference around four atomic program-evolution operators (Draft, Improve, Debug, Crossover): the same operators are trained via execution-grounded SFT and RL on data deduplicated against all evaluation benchmarks, then composed into long-horizon search, coupling learning and evolution in a single loop. On MLE-Bench Lite under a 12-hour per-task budget on one RTX 4090 capped at 12 GB VRAM, Frontis-MA1 (35B) improves Medal Average from 39.39% to 60.61% over its base model with OpenMLE-Evo, and reaches 71.21% with OpenMLE-Evo-Max (benchmark-independent experience priors and asynchronous search), exceeding GPT-5.5 + Codex and approaching GPT-5.6 Sol and the 2.8T Kimi K3. On held-out NatureBench Lite, both components transfer: with the framework fixed, swapping in the trained model raises Match-SOTA from 50% to 70%; with the model fixed, swapping in OpenMLE-Evo raises it from 20% to 50%. We release the model weights and the full OpenMLE stack to enable reproducible research on executable AI4AI toward RSI. Code: https://github.com/FrontisAI/OpenRSI
Problem

Research questions and friction points this paper is trying to address.

Recursive Self-Improvement
AI4AI
Machine Learning Engineering
Meta-evolution
Program Evolution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Recursive Self-Improvement
AI4AI
Machine Learning Engineering
Program Evolution
Execution-Grounded RL
🔎 Similar Papers
No similar papers found.
Junlin Yang
Junlin Yang
Department of Computer Science and Technology, Tsinghua University
Natural Language ProcessingMachine Learning
Che Jiang
Che Jiang
Tsinghua University
Y
Yu Fu
Horizon Research, Frontis.AI Tsinghua University
T
Tianwei Luo
Horizon Research, Frontis.AI Tsinghua University
C
Can Ren
Horizon Research, Frontis.AI Tsinghua University
Weizhi Wang
Weizhi Wang
Ph.D. Candidate, University of California Santa Barbara
Natural Language Processing
K
Kaikai Zhao
Horizon Research, Frontis.AI Tsinghua University
Hongyi Liu
Hongyi Liu
Zhili College, Tsinghua University
Machine learning
Y
Yuxin Zuo
Horizon Research, Frontis.AI Tsinghua University
Y
Yuru Wang
Horizon Research, Frontis.AI Tsinghua University
Yuchen Fan
Yuchen Fan
Shanghai AI Laboratory & Shanghai Jiao Tong University
NLPLarge Language ModelsEvaluation
K
Kai Tian
Horizon Research, Frontis.AI Tsinghua University
Z
Zhenzhao Yuan
Horizon Research, Frontis.AI Tsinghua University
X
Xiaojian Lin
Horizon Research, Frontis.AI Tsinghua University
L
Li Sheng
Horizon Research, Frontis.AI Tsinghua University
Rushi Qiang
Rushi Qiang
Tsinghua University, GaTech
Machine LearningAgents
G
Guoli Jia
Horizon Research, Frontis.AI Tsinghua University
Xingtai Lv
Xingtai Lv
Tsinghua University
Large Language ModelNatural Language Processing
Ermo Hua
Ermo Hua
Tsinghua University
Physics-driven Foundation Model
D
Dianqiao Lei
Horizon Research, Frontis.AI Tsinghua University
Youbang Sun
Youbang Sun
Assistant Researcher, Tsinghua University; Northeastern University; Texas A&M University
Distributed OptimizationMulti-Agent RLRiemannian OptimizationFederated Learning
Ning Ding
Ning Ding
Assistant Professor, Tsinghua University
Natural Language ProcessingMachine Learning
Bowen Zhou
Bowen Zhou
Chair Professor, Department of Electrical Engineering, Tsinghua University; Founder of Frontis.ai
Machine LearningNatural Language ProcessingRepresentation Learning and ReasoningConversational
Kaiyan Zhang
Kaiyan Zhang
Tsinghua University
Foundation ModelCollective IntelligenceScientific Intelligence