🤖 AI Summary
This study addresses the bottleneck of automatically pairing lightweight scene proxies with video data at scale by proposing a controllable world generation training paradigm that eliminates the need for paired data. Methodologically, the approach learns from ordinary RGBD videos, innovatively combining depth conditioning with joint RGBD generation. Furthermore, it incorporates cross-modal flow matching and a hybrid denoising strategy to effectively balance structural adherence with natural visual quality. To facilitate evaluation, a dedicated benchmark, ProxyBench, is constructed. Experimental results demonstrate that the proposed method outperforms existing baselines across multiple metrics, achieving significant improvements in both structural consistency and visual fidelity.
📝 Abstract
Lightweight scene proxies let creators control scene layout and motion while leaving room for imagination in appearance, lighting, and visual effects. However, a suitable proxy is not uniquely defined, making paired proxy-video data difficult to construct automatically at scale. We present Proxy2World, a controllable world model that learns these complementary capabilities from ordinary posed RGBD videos, without training on authored proxy-video pairs. The model jointly learns depth-conditioned RGB generation and joint RGBD generation through cross-modal flow matching. Learning both tasks enables proxy-camera hybrid denoising at inference to follow the proxy structure while producing natural, detailed visuals. We further introduce ProxyBench to evaluate this capability across a diverse set of scenes, camera trajectories, and subject motions. Experiments on ProxyBench show that Proxy2World achieves a better balance between structural adherence and visual quality than camera-controlled and geometry-conditioned methods, supported by quantitative metrics, VLM assessments, human evaluations and diverse qualitative results.