PhysEvo: Astra Can Act, Let It

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reliance of frozen models on external systems and their limited capacity for self-improvement in robotic manipulation by proposing PhysEvo, a framework that pioneers physical recursive self-improvement centered around a single frozen model. PhysEvo employs a dual-layer architecture comprising task agent execution and meta-agent iterative optimization, enabling the meta-agent to autonomously diagnose failures and revise tools and skills. This mechanism translates action consequences into persistent, testable system modifications without updating weights or training separate policies. Extensive evaluations on the RoboDojo simulation platform and AgileX PiPER real-world hardware demonstrate that PhysEvo achieves a 62% success rate across 42 simulated tasks and an average score of 90.6 with an 84% success rate in physical experiments, significantly outperforming baseline methods.
📝 Abstract
Astra can act, yet reliable manipulation depends on the system through which it observes and controls the world. We introduce PhysEvo, a framework for physical recursive self-improvement (RSI) around a single frozen model. A task agent executes robot tasks; a meta-agent uses the resulting trajectories to diagnose failures, revise tools and skills, and test corrections. The meta-agent can also improve its own diagnostic tools, so retained revisions support both later action and later self-improvement. This process develops joint-level control, evidence-seeking observation, and reusable manipulation skills without model-weight updates or a separately trained action policy. Across 42 RoboDojo tasks, held-out-layout evaluation of retained task-specific deployment versions yields a five-dimension average score of 68.14/100 and 62.00% success, compared with 47.17% for RoboDawn's one-shot Astra agent, the strongest published reference in our comparison. On eight manipulation tasks challenging direct Astra, PhysEvo achieves 55.00% success, compared with 1.25% for the direct-Astra reference. Deploying the simulation-evolved harness on AgileX PiPER and continuing skill revision yields 90.60/100 average score and 84.00% success across 25 trials on five real-world tasks. PhysEvo turns the consequences of action into persistent, testable changes to how a frozen model acts and improves.
Problem

Research questions and friction points this paper is trying to address.

physical recursive self-improvement
frozen model
robot manipulation
reliable control
skill revision
Innovation

Methods, ideas, or system contributions that make the work stand out.

Physical Recursive Self-Improvement
Frozen Model
Meta-Agent
Skill Revision
Sim-to-Real Transfer
🔎 Similar Papers
No similar papers found.