WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges faced by interactive video world models in long-horizon planning, where error accumulation often leads to failure and there is a lack of mechanisms to verify the long-term accuracy of arbitrary action sequences. The authors propose a self-verifiable reinforcement learning framework that models actions as consistent state operators rather than memorized temporal patterns by constructing invertible action cycles and repeatedly executing them. This approach introduces spatial closure and temporal consistency rewards, enabling unsupervised long-term state regression through the reversibility of action cycles. It supports generalization and verification of out-of-distribution compound action sequences. Experiments demonstrate up to a 44% reduction in state regression drift and nearly a fourfold improvement in compound action accuracy. The study also introduces CycleBench, a diagnostic benchmark for evaluating such capabilities.
📝 Abstract
Interactive video world models are essential for long-horizon planning and exploration, yet they suffer from compounding errors. Post-training methods such as reinforcement learning (RL) can improve these models, but they hit a verification bottleneck: for arbitrary action sequences, no ground-truth future state exists to measure long-term drift. Our key insight is that reversible action cycles make this verification possible: a sequence composed with its inverse must analytically return to the initial state, yielding annotation-free supervision on long-horizon correctness. Building on this, we introduce WorldCycle, a self-verifiable RL framework that constructs closed action cycles and their repeated executions from ordinary action sequences, and optimizes two complementary rewards: a spatial closure reward enforcing symmetry between mirrored forward and reverse segments, and a temporal consistency reward aligning states across repeated cycle executions. These rewards force the model to learn actions as consistent state operators rather than memorized temporal patterns, and extend naturally to out-of-distribution composite action cycles that the base model handles poorly. We further release CycleBench, a diagnostic benchmark for state-returning ability under complex action structures. WorldCycle reduces state returning drift by up to 44% and lifts composite-action accuracy nearly 4x over the base model, providing a vital foundation for physically grounded world models.
Problem

Research questions and friction points this paper is trying to address.

video world models
long-horizon planning
compounding errors
verification bottleneck
state drift
Innovation

Methods, ideas, or system contributions that make the work stand out.

self-verifiable reinforcement learning
video world models
action cycles
long-horizon planning
temporal consistency
🔎 Similar Papers
No similar papers found.
B
Bohai Gu
The Hong Kong University of Science and Technology
Y
Yueyang Yuan
Wuhan University
T
Taiyi Wu
AI Technology Center, Tencent Video, Tencent
Dazhao Du
Dazhao Du
Hong Kong University of Science and Technology
MultiModal LLMVideo UnderstandingTime Series ForecastingDeep Learning
J
Jian Liu
The Hong Kong University of Science and Technology
X
Xiaoyi Pang
The Hong Kong University of Science and Technology
J
Jie Zhang
The Hong Kong University of Science and Technology
X
Xiaocheng Lu
The Hong Kong University of Science and Technology
H
Haobin Zhong
AI Technology Center, Tencent Video, Tencent
X
Xiaotong Zhao
AI Technology Center, Tencent Video, Tencent
A
Alan Zhao
AI Technology Center, Tencent Video, Tencent
Song Guo
Song Guo
Chair Professor of CSE, HKUST
Large Language ModelEdge AIMachine Learning Systems