RoboFoundry: System-as-Policy Evolution for Self-Learning Embodied Agents

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of isolated component optimization and the difficulty of translating experience into persistent system improvements in embodied agents. To this end, it proposes the first self-evolving System-as-Policy framework, which treats support systems as unified policies subject to evolution. Through a co-evolutionary mechanism between contextual and hierarchical skill systems, along with shared semantic interface techniques, the framework enables capability transfer across heterogeneous robots and online autonomous evolution. Experimental results demonstrate that the proposed method achieves state-of-the-art performance on benchmarks such as EmbodiedBench while facilitating zero-shot transfer.
📝 Abstract
A foundation model should not act in isolation as an embodied agent. Yet, existing methods often optimize individual components of the agent stack, such as memory, context, skills, or action interfaces, rather than treating the supporting system itself as a unified policy. Moreover, interaction alone does not yield self-improvement unless execution experience is converted into persistent, validated system changes. We therefore propose RoboFoundry, the first embodied agentic framework that formulates this process as Self-Evolving System-as-Policy. RoboFoundry diagnoses capability gaps in decision-making and memory management, converts execution traces into validated task-specific system updates, and promotes recurring improvements to the general system. Evolution operates over two complementary surfaces: a context system that manages active internal context and persistent file-system memory, and a hierarchical skill system that organizes atomic skills, reusable compositions, and failure-conditioned recovery. A shared semantic interface separates embodiment-invariant decisions from embodiment-specific execution, allowing evolved system capabilities to transfer across heterogeneous robots. On EmbodiedBench, RoboFoundry achieves state-of-the-art performance, notably improving GPT-5.5 by 27.8%. It also brings Qwen3.7-Plus to near parity with GPT-5.5 (70.3% vs. 72.7%), showing consistent gains from system-as-policy evolution across foundation models. For long-horizon memory, RoboFoundry outperforms all baselines on RoboMemArena by at least 39.0%, even against methods assisted by external foundation models. On LIBERO-PRO, it further outperforms Cap-Agent0 by 243.8%-679.7% across all perturbation types. In real-world deployments, RoboFoundry demonstrates zero-shot transfer and online evolution across robots and tasks, highlighting its potential for fully autonomous embodied agents.
Problem

Research questions and friction points this paper is trying to address.

Embodied Agents
Self-Evolution
System-as-Policy
Foundation Models
Long-horizon Memory
Innovation

Methods, ideas, or system contributions that make the work stand out.

System-as-Policy
Self-Evolving
Embodied Agents
Hierarchical Skill System
Zero-Shot Transfer
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Jingsong Liang
Jingsong Liang
National University of Singapore
Robot LearningReinforcement LearningMulti-Agent SystemsEmbodied Intelligence
Shuhao Liao
Shuhao Liao
Beihang University
Multi-agent SystemsReinforcement LearningRobot learning
S
Shizhe Zhang
Nanyang Technological University
D
Diyuan Hou
Beihang University
Y
Yuxin Cai
Nanyang Technological University
X
Xinjian Deng
National University of Singapore
C
Chengyang He
National University of Singapore
W
Wenhui Huang
Nanyang Technological University
R
Runjia Tan
Nanyang Technological University
Z
Zhidong Wang
Nanyang Technological University
L
Lan Yu
Cloud Butterfly Technology
X
Xuesong Tian
Cloud Butterfly Technology
Guillaume Sartoretti
Guillaume Sartoretti
Assistant Professor, National University of Singapore (NUS), Mechanical Engineering Dpt
Multi-Agent SystemsRoboticsSwarm IntelligenceDistributed ControlDistributed Learning
J
Jie Luo
Beihang University
Y
Yao Mu
Shanghai Jiao Tong University
W
Wenjun Wu
Beihang University
Wanhua Li
Wanhua Li
Harvard University
Computer VisionPattern Recognition
C
Chen Lv
Nanyang Technological University