NAMVIS: Next-Scale Autoregressive Multi-View Image Synthesis

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high inference costs and limited multi-view generation efficiency of diffusion models in sparse-view synthesis by proposing a diffusion-free framework that reformulates multi-view synthesis as geometry-conditioned next-scale autoregressive modeling. Methodologically, it introduces multi-scale projection pose encoding to inject camera transformations into the attention mechanism, combines global conditioning with dense geometry-aware cross-attention to ensure multi-view consistency, and enables parallel target view generation through coarse-to-fine discrete visual token prediction. Experimental results demonstrate that the proposed method outperforms diffusion-based baselines in both PSNR and SSIM on benchmarks such as Objaverse while achieving over three-fold faster inference, validating its potential as an efficient alternative for sparse-view synthesis.
📝 Abstract
Sparse-view novel view synthesis is a central problem in 3D content creation, but diffusion-based approaches remain limited by iterative denoising, making multi-view generation expensive at inference time. We introduce NAMVIS, a diffusion-free framework that reformulates multi-view image synthesis as geometry-conditioned next-scale autoregression. Instead of generating target views through repeated denoising, NAMVIS predicts discrete visual tokens through a small number of coarse-to-fine scale steps, while sampling all tokens within each scale and across target views in parallel. To anchor this generation process to explicit camera geometry, we propose Multi-scale Projective Pose Encoding, which injects source and target camera transformations into both target-view self-attention and source-to-target cross-attention at every resolution. NAMVIS further combines global conditioning with dense geometry-aware cross-attention, enabling the model to preserve source-view appearance while maintaining target-view consistency. Across Objaverse, GSO, and OmniObject3D, NAMVIS outperforms diffusion-based baselines in PSNR, SSIM, and LPIPS, while running over 3 times faster than the evaluated diffusion baselines under the same evaluation setting. These results suggest that geometry-conditioned next-scale autoregression is a promising and efficient alternative to diffusion for sparse-view multi-view synthesis. Additional qualitative results, videos, and resources are available at https://corl-team.github.io/namvis/
Problem

Research questions and friction points this paper is trying to address.

sparse-view novel view synthesis
multi-view image generation
diffusion models
inference efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Next-Scale Autoregression
Multi-View Image Synthesis
Diffusion-Free
Multi-scale Projective Pose Encoding
Sparse-View