🤖 AI Summary
This study addresses the long-standing separation between reconstruction and generation tasks in existing visual models, which often leads to blurry outputs or geometric inconsistencies. To bridge this gap, it proposes a unified flow-based framework demonstrating that both tasks can share mathematical formulations and backbone architectures, with behavioral differences governed solely by denoising configurations. Specifically, through a shared clean target predictor, single-step endpoint prediction performs reconstruction while multi-step flow unrolling enables conditional generation. This paradigm significantly simplifies model unification. Experimental results show that JiT-LVSM improves novel view synthesis quality and JUSt3R effectively reduces geometric artifacts, thoroughly validating the efficacy of the proposed unified framework.
📝 Abstract
Reconstruct where the images provide evidence, and generate where they do not: recent success of spatial world models such as Atlas (World Labs Team, 2026) highlights the value of unifying reconstruction and generation in one model. Yet the two have long lived in separate paradigms with distinctive failure modes: feed-forward reconstruction averages ambiguity into blur, while conditional generation invents plausible but scene-inconsistent detail. In this work, we present a unified flow-based formulation for reconstruction and generation, where a shared clean-target predictor performs direct reconstruction at its single-step endpoint and unfolds conditional generation through multi-step flow. A controlled toy study reveals the mechanism: with a single step, the predictor collapses to the conditional mean just like feed-forward methods, favoring consistency over diversity. With multi-step inference, the fidelity of generated details grows with context richness: closer observations reduce ambiguity and yield better-matched details. We further instantiate the formulation in appearance and geometry 3D tasks. JiT-LVSM improves perceptual and distributional quality in novel view synthesis, while JUSt3R retains competitive single-step geometry prediction with additional multi-step inference capabilities that reduces veil and flying-pixel artifacts, producing cleaner surface structure with greater test-time compute. Together, they show that reconstruction and generation can share both a formulation and a backbone, with their behavior governed by denoising configuration---making unification surprisingly simple.