From Feed-Forward to Flow: Unifying Reconstruction and Generation Is Easier Than You Think

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the long-standing separation between reconstruction and generation tasks in existing visual models, which often leads to blurry outputs or geometric inconsistencies. To bridge this gap, it proposes a unified flow-based framework demonstrating that both tasks can share mathematical formulations and backbone architectures, with behavioral differences governed solely by denoising configurations. Specifically, through a shared clean target predictor, single-step endpoint prediction performs reconstruction while multi-step flow unrolling enables conditional generation. This paradigm significantly simplifies model unification. Experimental results show that JiT-LVSM improves novel view synthesis quality and JUSt3R effectively reduces geometric artifacts, thoroughly validating the efficacy of the proposed unified framework.
📝 Abstract
Reconstruct where the images provide evidence, and generate where they do not: recent success of spatial world models such as Atlas (World Labs Team, 2026) highlights the value of unifying reconstruction and generation in one model. Yet the two have long lived in separate paradigms with distinctive failure modes: feed-forward reconstruction averages ambiguity into blur, while conditional generation invents plausible but scene-inconsistent detail. In this work, we present a unified flow-based formulation for reconstruction and generation, where a shared clean-target predictor performs direct reconstruction at its single-step endpoint and unfolds conditional generation through multi-step flow. A controlled toy study reveals the mechanism: with a single step, the predictor collapses to the conditional mean just like feed-forward methods, favoring consistency over diversity. With multi-step inference, the fidelity of generated details grows with context richness: closer observations reduce ambiguity and yield better-matched details. We further instantiate the formulation in appearance and geometry 3D tasks. JiT-LVSM improves perceptual and distributional quality in novel view synthesis, while JUSt3R retains competitive single-step geometry prediction with additional multi-step inference capabilities that reduces veil and flying-pixel artifacts, producing cleaner surface structure with greater test-time compute. Together, they show that reconstruction and generation can share both a formulation and a backbone, with their behavior governed by denoising configuration---making unification surprisingly simple.
Problem

Research questions and friction points this paper is trying to address.

reconstruction
generation
unification
spatial world models
feed-forward
Innovation

Methods, ideas, or system contributions that make the work stand out.

Flow-based formulation
Unified reconstruction and generation
Shared clean-target predictor
Multi-step inference
Novel view synthesis
🔎 Similar Papers