🤖 AI Summary
This study addresses the challenge that future predictions from generative world models are difficult to translate into compact control targets, leaving terminal precision without explicit supervision. We propose an entity-level target readout interface that, for the first time, treats the prediction-to-execution bridge as an explicitly learned component, converting 3D trajectory world model predictions into executable SE(3) targets. This approach integrates object-centric pose prediction with depth-grounded translation to optimize target representation. Closed-loop control is achieved through a shared pose-native executor combined with online object pose feedback. Experimental results demonstrate that the proposed method attains an average success rate of 79.69% across five manipulation tasks and achieves up to 75% zero-shot deployment success on a Franka robotic arm.
📝 Abstract
Generative world models provide rich predictions of how manipulation scenes may evolve toward task objectives, yet those futures do not directly expose the compact task variables required by control. When training supervises future prediction alone, terminal goal accuracy is not an explicit learning objective, even when geometric recovery is available. We present Entity-Level Goal Readout, a learned prediction-to-execution interface that makes the executable terminal goal an explicit output of a 3D trace world model. It combines object-centric pose prediction with translation grounded in observed depth to produce a compact goal in SE(3). A shared Pose-Native Executor consumes this fixed goal with online object-pose feedback for closed-loop control without rerunning the world model. Across five manipulation tasks, the pipeline achieves a mean success rate of 79.69%. Goal diagnostics directly measure terminal goal accuracy, while controlled translation perturbations characterize how execution degrades under goal error. Zero-shot deployment on a Franka arm achieves 73.33% success on nominal StackCube, 66.67% with distractors, and 75.00% on PickPlate with a target unseen during policy training. These results support treating the prediction-to-execution interface as an explicit learned component of world-model planning rather than incidental post-processing in the control pipeline itself. Project page: https://claire0730.github.io/executable-goals/