HL-OutPaint: Coarse-to-Fine Video Outpainting for High-Resolution Long-Range Videos

πŸ“… Unknown Date
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Existing video outpainting methods struggle to simultaneously achieve large-scale spatial extension and long-term global spatiotemporal consistency. This work proposes a coarse-to-fine two-stage generation framework: first, a novel global–local frame swapping mechanism constructs a low-resolution Global Coarse Guidance (GCG) that unifies long-range structural consistency with short-term dynamics within a shared representation; subsequently, high-resolution details are synthesized under the guidance of the GCG. By effectively decoupling global modeling from local generation, the proposed method significantly outperforms existing approaches in both large-scale spatial outpainting and long-duration video synthesis, producing results that exhibit rich detail while maintaining strong spatiotemporal coherence.
πŸ“ Abstract
Video outpainting generates plausible visual content beyond the original spatial extent of a video, playing a key role in adapting videos to diverse display formats. To support such use cases, it must enable large spatial extrapolation over long sequences. However, most existing methods address only one of these challenges or lack explicit mechanisms for ensuring global spatio-temporal consistency, leading to notable limitations. In this paper, we propose HL-OutPaint, a high-resolution video outpainting framework for long sequences. Our approach follows a coarse-to-fine strategy with a two-stage pipeline. We first construct Global Coarse Guidance (GCG), a low-resolution representation that captures global structure and dominant motion across the video. Unlike naive downsampling, GCG is built via a novel global-local frame swapping mechanism that couples sparse global keyframes with local temporal windows and exchanges information during sampling. This enables GCG to encode both long-term structural consistency and short-term temporal dynamics in a unified representation. Guided by this representation, HL-OutPaint then performs high-resolution outpainting to generate spatially detailed and temporally consistent content. By separating global structure modeling from fine-grained synthesis, our framework achieves stable, coherent generation for large spatial expansion and long video sequences. Extensive experiments show that HL-OutPaint outperforms existing methods in challenging scenarios involving wide spatial extrapolation and long video sequences.
Problem

Research questions and friction points this paper is trying to address.

video outpainting
spatio-temporal consistency
long-range video
high-resolution
spatial extrapolation
Innovation

Methods, ideas, or system contributions that make the work stand out.

video outpainting
coarse-to-fine
global-local frame swapping
spatio-temporal consistency
high-resolution video generation
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Jeongeun Park
Jeongeun Park
Ph D. Student, Korea University
roboticsartificial intelligence
Janghyeok Han
Janghyeok Han
M.S. student at POSTECH
Video restorationImage restoration
Geonung Kim
Geonung Kim
Ph.D Student, POSTECH
Computational PhotographyDeep Learning
H
Hyun-Seung Lee
Visual Display Business, Samsung Electronics, Republic of Korea
K
Kyuha Choi
Visual Display Business, Samsung Electronics, Republic of Korea
Y
Youngseok Han
Visual Display Business, Samsung Electronics, Republic of Korea
Sunghyun Cho
Sunghyun Cho
POSTECH
Computer GraphicsComputer VisionImage ProcessingComputational Photography