🤖 AI Summary
This study addresses the performance degradation of vision-language-action models on out-of-distribution tasks by proposing an inference-time internal representation steering method. Without updating policy parameters, the approach trains a linear classifier solely on successful and failed trajectories to identify intervention points and derive steering directions, enabling effective intervention on both the frozen policy π0.5 and the world action model Cosmos Policy. Experimental results demonstrate that the success rate in simulated tasks improves from 44.4% to 66.2%, while real-world robotic manipulation performance is also substantially enhanced. These findings validate the effectiveness and generality of this retraining-free paradigm in restoring model generalization capabilities.
📝 Abstract
Vision-language-action (VLA) and world-action models (WAMs) often degrade under out-of-distribution task variations despite retaining partial task capability. To recover such capability, we propose RoboIRS, an inference-time internal representation steering method that uses successful and failed rollouts to train linear classifiers, select outcome-relevant intervention locations, and derive task-specific steering directions without updating policy parameters. On 15 simulation tasks with a frozen $\pi 0.5$ policy, RoboIRS improves the average success rate from 44.4% to 66.2%, outperforming alternative inference-time intervention baselines while adding little inference time. We further validate RoboIRS on real-robot manipulation using the same $\pi 0.5$ policy and demonstrate its applicability to a world-action model Cosmos Policy, where the average success rate improves from 35.4% to 55.4%. These results show that directly steering internal robot-policy representations can improve the performance of robot policies at inference time. Project website is available at https://rollingoat.github.io/roboirs/.