🤖 AI Summary
This study addresses the limitation that failure data are typically discarded during vision-language-action policy training, leaving robots incapable of anomaly recovery. To overcome this, we propose a simulation-based method for transforming failed trajectories into valuable training signals. The core innovation lies in defining the concept of intervention recoverability, which combines adaptive Monte Carlo estimation with Wilson score intervals to precisely identify frontier states exhibiting low recoverability. Leveraging the SmolVLA architecture, these identified states enable the conversion of failure trajectories into high-value synthetic recovery data. Evaluated on Franka manipulation tasks, the proposed approach achieves a recovery success rate of 38.4%, outperforming uniform sampling by 6.7 percentage points. Furthermore, it maintains a significant advantage under perturbed environmental conditions, effectively validating its robustness.
📝 Abstract
Simulation enables scalable training of Vision-Language-Action policies by using privileged experts to generate visual demonstrations without requiring every trajectory to be collected through manual teleoperation. However, such pipelines typically retain successful demonstrations while failed rollouts are discarded, even though they expose precisely the off-nominal states from which recovery must be learned.
We introduce Kintsugi-VLA, a framework for converting failed rollouts into targeted synthetic recovery data by exploiting exact state restoration and branching in simulation. For a fixed privileged expert, we define interventional recoverability as the probability of completing the original task after the simulator is restored to a given state, estimate it using adaptive Monte Carlo continuations with pointwise Wilson intervals, and characterize its non-monotonic evolution along failed trajectories. These estimates identify an observed terminal low-recoverability frontier-the point after which measured recoverability remains below a threshold-which is then used to select informative recovery starting states.
In a simulated Franka manipulation task, targeted recovery data yield aggregate SmolVLA recovery success of 34.6\% and 38.4\% under difficulty- and frame-budget matching, respectively, 5.8 and 6.7 percentage points above uniform sampling within the same recovery window. The same ordering is observed under disturbed end-to-end execution and shifted clutter and physics conditions, while clean-task success decreases from 76.8\% to 74.7\%. Kintsugi-VLA demonstrates how failed simulator rollouts can be transformed from discarded experience into structured recovery-training data through direct interventional measurement.