🤖 AI Summary
This study addresses the problem that temporal misalignment in generated videos can render robotic manipulation guidance ineffective or even detrimental. To overcome this, we propose a reliability-aware future conditioning method that pioneers formulating temporal alignment as a control problem. By integrating digital twins, unmasked video diffusion models, and candidate ensembling, our framework constructs a future-experience-conditioned paradigm that dynamically evaluates and selects trustworthy future hypotheses using only task rewards, without requiring alignment labels or supervision signals. Experiments demonstrate that this approach significantly mitigates the adverse effects of temporal mismatch, improving real-world robotic manipulation success rates from 26.7% to 56.7% and substantially outperforming existing baseline strategies.
📝 Abstract
A generated video of a task the robot is about to perform is useful guidance only if it depicts the phase the robot is actually in. We show that temporal misalignment can turn a task-consistent generated future into actively harmful guidance. On CALVIN, a five-frame early shift nearly erases the benefit of generated futures, reducing success from 81.3% to 54.8% against 54.0% without futures; imposed timing shifts reduce it even further to 34.2%, 19.8 points below the future-free policy. We introduce Reliability-Aware Future Conditioning (RAFC), which treats this as a control problem rather than a generation problem. At every step, RAFC estimates how far to trust the received clip and which nearby temporal hypothesis to prefer, falling back toward a static branch when neither fits, and it learns both from task reward alone without shift labels or alignment supervision. RAFC sits on top of Future-Experience Conditioning (FEC), which builds the clip once from task grounding, a robot-free digital-twin rollout, and mask-free video diffusion. Under deliberately off-grid phase shifts and rate mismatch, RAFC substantially improves success under temporal mismatch. Candidate ensembling accounts for most of the recovery near alignment, while learned reliability adds a further 7.0 percentage points over uniform averaging of the identical candidate bank under off-grid shifts. The gain holds on the evaluated task sets and survives on a Franka under natural timing mismatch nobody imposed, where aggregate success rises from 26.7% to 56.7%. All resources will be made publicly available. https://future-condition.github.io/.