LIBERO-MAX: Do Robot Policies Adapt When the World Changes?

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the difficulty of robotic policies in adapting to dynamic environmental changes during task execution and the lack of standardized evaluation. To this end, we introduce LIBERO-MAX, a simulation benchmark featuring a novel controlled paired protocol that isolates the impact of mid-episode events, effectively distinguishing event-induced regressions from inherent failures. This framework systematically evaluates the robustness of Vision-Language-Action (VLA) models and other policies against geometric and observational perturbations. Our analysis reveals shared vulnerabilities across policies and trajectory-dependent effects, quantitatively demonstrating that mid-episode events degrade success rates by 11.0 to 25.7 percentage points. Ultimately, this work establishes a reproducible diagnostic and evaluation platform for analyzing failure modes in embodied intelligence policies.
📝 Abstract
Robots must often continue a task after a target moves, the viewpoint shifts, or an obstacle appears, even though their earlier observations and committed actions reflect the previous scene. Many simulation robustness benchmarks fix external conditions at reset, leaving this temporal challenge underexamined. We introduce LIBERO-MAX, a benchmark of 8,000 paired cases spanning eight types of changes to geometry, observations, appearance, clutter, and paths. Each pair compares task execution with and without a mid-task event, holding the task, initial state, policy seed, and pre-event action sequence fixed. This controlled comparison distinguishes event-associated regressions from failures already present without the change. Across fourteen current VLA, hybrid, and world-action policies, events reduce success by 11.0-25.7 percentage points. Event profiles reveal shared vulnerabilities to geometry and observation changes, while policy-family rankings interleave. Camera controls show that robustness reflects both competence under the changed conditions and the trajectory from which they are encountered; varying query cadence does not eliminate the gap. Together, the paired protocol and temporal diagnostics establish LIBERO-MAX as a reproducible testbed for diagnosing failures under mid-execution changes and measuring progress toward robot policies that remain effective as the world changes.
Problem

Research questions and friction points this paper is trying to address.

robot policy adaptation
mid-task environmental changes
robustness benchmark
temporal challenge
Innovation

Methods, ideas, or system contributions that make the work stand out.

mid-task robustness
paired evaluation protocol
temporal diagnostics
robot policy benchmark
environmental perturbation