ASENA: Self-evolving Agents for Embodied Navigation

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of enabling embodied agents to self-evolve and navigate efficiently under fixed model weights. To this end, it proposes a persistent experience evolution mechanism that operates without weight updates, establishing a system bridging code-based agents with robotic perception to support program generation, execution repair, and skill reuse. Furthermore, the method integrates a 4B-parameter monocular vision-language navigation strategy, achieving closed-loop optimization through a shared decoder and a geometric atomic task dataset. Experimental results demonstrate a 98% success rate on the R2R benchmark with significantly reduced interaction steps, while real-world evaluations confirm successful execution of complex behaviors in unmapped environments.
📝 Abstract
We present ASENA, an embodied agent system that connects general-purpose coding agents to robot sensing, computation, supervised execution, and persistent experience. Agents can write and execute programs, inspect recorded outcomes, repair failures, and reuse notes and executable skills while keeping their model weights fixed. We further introduce ASENA-VLN, a 4B monocular navigation policy that serves as an optional tool within this programmable system. ASENA-VLN predicts body-frame trajectories for both extended routes and short-horizon behaviors using a shared vision-language decoder trained on route instructions, visual question answering, and a newly curated dataset of geometry-derived atomic navigation tasks. As a standalone policy, ASENA-VLN achieves state-of-the-art success rates of 68.7% on R2R and 70.2% on RxR. When integrated with a coding agent, learned navigation improves ASENA's success rate by 11 percentage points on both agentic benchmarks while reducing execution time. Through persistent workspace evolution and simulator feedback, ten passes over recurring 100-task subsets further improve success from 72% to 98% on R2R and from 65% to 89% on RxR. On embodied question answering, ASENA achieves state-of-the-art accuracy with fewer interaction steps. Finally, real-world demonstrations on a Unitree G1 combine search, visual inspection, spatial reasoning, and synthesized gestures without a pre-built map, illustrating how online programming extends robot behavior beyond route following and predefined skills.
Problem

Research questions and friction points this paper is trying to address.

Embodied Navigation
Self-evolving Agents
Vision-Language Navigation
Embodied Question Answering
Programmable Agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

Embodied Navigation
Self-evolving Agents
Vision-Language Navigation
Programmable Agents
Persistent Experience
🔎 Similar Papers
No similar papers found.