ReSync: Re-Aligning the Two Clocks of Asynchronous World-Action Models

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the asynchronous clock mismatch between video and action streams in world models, where future predictions fail to converge before actions are executed. To tackle this, we formalize a dual-clock mechanism and introduce a commitment-evidence gap metric, quantified via diffusion denoising scheduling. We propose advancing the world model exclusively within a support window while maintaining action states to achieve resynchronization, requiring neither parameter modifications nor candidate comparisons. The primary contribution is an efficient, zero-parameter-tuning asynchronous inference optimization strategy. Evaluated on the RoboCasa simulation benchmark, our approach improves success rates by 4.48 percentage points over a wait-control baseline under equivalent computational budgets, with the proposed rules demonstrating transferability across other benchmarks and backbone architectures.
📝 Abstract
Jointly generating future video and actions has become a standard recipe for world-action models, and the strongest systems denoise the two streams on separate schedules: actions are decoded in few steps so control stays fast, while the video stream runs longer to keep the predicted future sharp. The design is deliberate, but it leaves the two streams on different clocks, and an action can become executable while the future that should justify it is still largely unresolved. We formalize this as a two-clock view of asynchronous inference and introduce the commitment-evidence gap, a quantity read directly from a model's own sampling schedule rather than measured by search. The gap is predictive: as it widens, candidate utility becomes harder to identify and extra candidate sampling buys less, while advancing the world stream buys more, and the two cross. Spending more world computation is therefore not simply better. The useful interval is closed at both ends, and both ends can be read off the schedule before any rollout. ReSync places the computation inside it: hold the action state, advance only the world within the supported window, then resume native denoising. No parameters change and no candidates are compared. On a frozen paired RoboCasa panel this improves success by 4.48 points, while an equal-compute control that waits without advancing the world does not move, and the same rule transfers to a second benchmark and a second backbone without retuning.
Problem

Research questions and friction points this paper is trying to address.

world-action models
asynchronous inference
commitment-evidence gap
two-clock misalignment
video-action generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

asynchronous world-action models
commitment-evidence gap
two-clock inference
ReSync
training-free alignment
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xi Lin
Johns Hopkins University
F
Feihong Zhang
Tsinghua University
Y
Yulong Shi
Yinwang Intelligent Technology Co., Ltd.
Y
Yanghong Mei
University of Chinese Academy of Sciences
Z
Zuxing Lu
Yinwang Intelligent Technology Co., Ltd.
X
Xiaofan Zhu
Yinwang Intelligent Technology Co., Ltd.
Z
Zihao Liang
Yinwang Intelligent Technology Co., Ltd.
Zhirui Gao
Zhirui Gao
National University of Defense Technology
Computer vision3D reconstructionDifferentiable rendering
Zhaowen Li
Zhaowen Li
National Laboratory of Pattern Recognition,Institute of Automation,Chinese Academy of Sciences
Computer VisionArtificial IntelligenceSelf-supervised Learning