LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of synthesizing first-person videos from exocentric views, where limited view overlap leads to depth estimation errors and content misalignment. To overcome these issues, this work proposes a lift-free framework that eliminates explicit 3D reconstruction and reprojection. Specifically, it leverages an LVLM-style Transformer to learn cross-view probabilistic mappings for structural alignment, while integrating video diffusion models with confidence-guided denoising to restore fine-grained texture details. The proposed approach outperforms existing state-of-the-art pipelines relying on explicit 3D representations and demonstrates zero-shot generalization across datasets without additional training. Ultimately, this method effectively balances structural integrity with visual fidelity in cross-view video generation.
📝 Abstract
Generating an egocentric video from a single exocentric recording is a challenging case of novel view synthesis, as the two cameras share little overlap and much of the target view is unobserved. Current state-of-the-art methods reconstruct the scene explicitly by estimating depth, lifting the video into a point cloud, and re-rendering it from the egocentric camera to condition a video diffusion model. This deterministic mapping assigns each pixel to a single reprojected location, which preserves texture but translates depth errors into misplaced content. We ask what a video diffusion model should receive as its condition and propose a lifting-free answer: a learned view synthesizer, an LVSM-style transformer fine-tuned to render the egocentric view directly without depth, point clouds, or reprojection, resolving cross-view correspondence internally. In contrast, its probabilistic mapping averages each region over candidate source locations according to a learned correspondence distribution, preserving structure while fine texture is averaged away. We argue that this trade-off suits a diffusion generator, whose denoising training excels at restoring detail, so an effective condition should prioritize structural alignment over sharpness. This distribution's concentration also yields a per-region confidence, used both to mask low-confidence regions and to guide the generator toward high-confidence areas during early layout-forming denoising steps. Our approach consistently outperforms the state-of-the-art explicit pipeline and generalizes to other datasets without retraining. The synthesizer thus supplies view structure, and the diffusion model its detail.
Problem

Research questions and friction points this paper is trying to address.

Exocentric-to-Egocentric Video Generation
Novel View Synthesis
Video Diffusion Model
Cross-view Correspondence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Lifting-Free View Synthesis
Exocentric-to-Egocentric Video Generation
Video Diffusion Model
Probabilistic Mapping
Learned View Synthesizer
🔎 Similar Papers
No similar papers found.