ProactiveVLA: Augmenting Embodied Memory through Proactive Environment Exploration

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation that robots struggle to acquire comprehensive interaction knowledge in novel environments through mere task repetition. To overcome this, it proposes an active exploration strategy integrating a reasoning agent with a frozen vision-language-action (VLA) model. By leveraging object affordances, state changes, and compositional interactions, the system autonomously generates exploration goals. Execution feedback and experiences from self-proposed objectives are then consolidated into reusable embodied memory, guiding subsequent planning and control while transcending the constraints of passive task refinement. Experimental results demonstrate that this approach outperforms existing baselines on LIBERO-Pro and RoboCasa365. Notably, under stringent constraints, it achieves a 48% success rate on the LIBERO-Pro Goal-T benchmark, substantially surpassing the state-of-the-art performance of 19%.
📝 Abstract
Rapid adaptation to a new environment requires a robot to acquire useful knowledge about local objects, states, and interactions from limited experience. Systems that combine a reasoning agent with a frozen vision-language-action model (VLA) can adapt through execution feedback and memory, making the choice of experience central to their effectiveness. Repeated practice of a target task may refine a familiar solution while leaving other interactions relevant to changed conditions untested. We introduce ProactiveVLA, which uses proactive environment exploration to acquire reusable knowledge for deployment-time adaptation. After completing an initial task, the agent allocates the remaining interaction budget to self-proposed goals covering object affordances, state-changing interactions, and compositions of interactions. It verifies execution outcomes and consolidates both task-directed and exploratory experience into memory that guides subsequent planning and control. ProactiveVLA outperforms the baselines under the same turn budget on LIBERO-Pro and RoboCasa365 Composite-Seen. On LIBERO-Pro Goal-T, with at most one VLA primitive invocation allowed during evaluation, ProactiveVLA completes 48% of instances, compared with 19% for the state-of-the-art task-refinement baseline.
Problem

Research questions and friction points this paper is trying to address.

embodied AI
vision-language-action model
environment exploration
rapid adaptation
robot memory
Innovation

Methods, ideas, or system contributions that make the work stand out.

Proactive Exploration
Vision-Language-Action Model
Embodied Memory
Deployment-time Adaptation
Affordance Discovery
🔎 Similar Papers
2024-10-04International Conference on Learning RepresentationsCitations: 0