🤖 AI Summary
This study addresses the reliance on external annotations and error propagation from geometry-guided attention in third-person to first-person video generation by proposing a text-free, geometry-bias-free framework. Methodologically, it introduces a dynamic captioning mechanism that extracts conditioning information directly from diffusion model hidden states, implicitly learning cross-view correspondences to replace external text supervision. Additionally, it integrates lightweight depth priors with large-scale multi-view joint training, eliminating the need for computationally expensive geometry-guided attention. The proposed framework achieves state-of-the-art performance on the Ego-Exo4D benchmark. Notably, it requires no external annotations during inference while significantly accelerating generation speed, demonstrating superior generalization capabilities in unconstrained in-the-wild scenarios.
📝 Abstract
Exo-to-egocentric video generation aims to synthesize what a person sees from their own viewpoint given third-person footage and a target head trajectory. The task requires transferring appearance and semantics across large viewpoint changes while hallucinating content never observed by the exocentric camera. Existing approaches either impose additional input requirements, such as a ground-truth initial egocentric frame or multiple synchronized exocentric views, or remain limited to category-specific settings. EgoX is the first to address cross-activity and in-the-wild generalization, but requires a human-provided caption of the non-existent egocentric view at inference and introduces a computationally expensive geometry-guided attention bias that can propagate reconstruction errors and suppress textual and visual context. We therefore propose \textbf{Ego-Forge}, a caption-free and bias-free framework for exo-to-egocentric generation. It introduces \textit{Dynamic Captioning}, which derives conditioning tokens directly from the model's hidden states and adapts them to the diffusion timestep and network depth, replacing external text conditioning. By scaling training by an order of magnitude and using all available exocentric viewpoints, Ego-Forge learns cross-view correspondence implicitly and eliminates the need for geometry-guided attention, requiring only a lightweight depth prior. Ego-Forge achieves state-of-the-art performance on Ego-Exo4D, runs faster end-to-end, requires no external annotation at inference, and generalizes to in-the-wild scenes, including cases where over-reliance on geometry blocks appearance inference. Our model and source code will be made publicly available.