MotionCraft: Latent World Modeling with Sparse Attention for Visual Upscaling

📅 2026-08-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge in video super-resolution of simultaneously preserving fine local details, modeling long-range spatiotemporal dependencies, achieving perceptual realism, and maintaining computational efficiency. To this end, the authors propose a motion-aware latent state prediction framework featuring several key innovations: a motion-informed latent world representation, a Latent World Transformer that balances local and non-local interactions, an adaptive sparse attention mechanism, and a compact conditional decoder enabling user-controllable reconstruction. The proposed method achieves significant improvements in both reconstruction accuracy and perceptual quality while retaining high computational efficiency. Moreover, it offers a predictable trade-off between temporal smoothness and detail fidelity, allowing users to tailor the output according to specific application needs.
📝 Abstract
Video super-resolution (VSR) aims to recover high-fidelity high-resolution videos from low-resolution inputs and is central to applications ranging from mobile capture to streaming and archival restoration. Existing approaches trade off among local-detail fidelity, long-range spatio-temporal modeling, perceptual realism, and efficiency: convolutional alignment techniques preserve local structure but suffer when motion is large or degradations are complex; transformer-based methods capture long-range dependencies yet require architectural or algorithmic adaptations to remain computationally feasible; and recent latent or diffusion-based generators synthesize rich texture but require specialized temporal constraints to maintain coherence. We present MotionCraft, a controllable VSR framework that formulates restoration as motion-aware latent state prediction inspired by world models and integrates adaptive sparse attention with an explicit user-accessible control interface. MotionCraft combines robust motion fusion, a Latent World Transformer that balances locality and targeted non-local interactions, and a compact conditional decoder to deliver temporally consistent, high-quality reconstructions under streaming constraints. Empirical evaluations show that MotionCraft achieves strong reconstruction and perceptual performance while enabling predictable trade-offs between temporal smoothness and reconstruction fidelity.
Problem

Research questions and friction points this paper is trying to address.

video super-resolution
spatio-temporal modeling
perceptual realism
computational efficiency
temporal coherence
Innovation

Methods, ideas, or system contributions that make the work stand out.

latent world modeling
sparse attention
motion-aware super-resolution
controllable video restoration
temporal consistency
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Rong Fu
Independent Researcher
Chunlei Meng
Chunlei Meng
Fudan University
Embodied Ai,Multimodal,Multi-agent
Y
Yangchen Zeng
Independent Researcher
Xiaowen Ma
Xiaowen Ma
Zhejiang University, Huawei Noah's Ark Lab
Computer VisionRemote SensingMulti-modalTime Series
Y
Yongtai Liu
Independent Researcher
Wangyu Wu
Wangyu Wu
The University of Liverpool
S
Shuo Yin
Independent Researcher
Z
Zijian Zhang
Independent Researcher
S
Sicheng Li
Independent Researcher
Y
Yingrui Ji
Independent Researcher
Chenhao Wang
Chenhao Wang
Tencent
Natural Language ProcessingLarge Language Models
Simon Fong
Simon Fong
Associate Professor, University of Macau
Data Mining and Optimization