🤖 AI Summary
This study addresses the asynchronous clock mismatch between video and action streams in world models, where future predictions fail to converge before actions are executed. To tackle this, we formalize a dual-clock mechanism and introduce a commitment-evidence gap metric, quantified via diffusion denoising scheduling. We propose advancing the world model exclusively within a support window while maintaining action states to achieve resynchronization, requiring neither parameter modifications nor candidate comparisons. The primary contribution is an efficient, zero-parameter-tuning asynchronous inference optimization strategy. Evaluated on the RoboCasa simulation benchmark, our approach improves success rates by 4.48 percentage points over a wait-control baseline under equivalent computational budgets, with the proposed rules demonstrating transferability across other benchmarks and backbone architectures.
📝 Abstract
Jointly generating future video and actions has become a standard recipe for world-action models, and the strongest systems denoise the two streams on separate schedules: actions are decoded in few steps so control stays fast, while the video stream runs longer to keep the predicted future sharp. The design is deliberate, but it leaves the two streams on different clocks, and an action can become executable while the future that should justify it is still largely unresolved. We formalize this as a two-clock view of asynchronous inference and introduce the commitment-evidence gap, a quantity read directly from a model's own sampling schedule rather than measured by search. The gap is predictive: as it widens, candidate utility becomes harder to identify and extra candidate sampling buys less, while advancing the world stream buys more, and the two cross. Spending more world computation is therefore not simply better. The useful interval is closed at both ends, and both ends can be read off the schedule before any rollout. ReSync places the computation inside it: hold the action state, advance only the world within the supported window, then resume native denoising. No parameters change and no candidates are compared. On a frozen paired RoboCasa panel this improves success by 4.48 points, while an equal-compute control that waits without advancing the world does not move, and the same rule transfers to a second benchmark and a second backbone without retuning.