🤖 AI Summary
This work addresses the scalability bottleneck in feed-forward novel view synthesis, where Transformer-based architectures suffer from prohibitive computational and memory overheads as the number of input views increases. To this end, we propose the AIMS framework, whose core innovation lies in decoupling the number of observations from the number of processed views via an anchor integration mechanism. Specifically, AIMS employs farthest point sampling to select a fixed set of anchors, combined with spatial clustering and a lightweight learnable integrator to aggregate information from neighboring observations. This design enables the efficient utilization of large-scale multi-view data under a constant global view budget. Experimental results demonstrate that AIMS achieves PSNR values of 29.41 dB on RealEstate10K and 17.73 dB on ScanNet, with an average rendering time of only 7.24 milliseconds per view, yielding a superior quality-efficiency trade-off.
📝 Abstract
Feed-forward novel view synthesis methods achieve strong generalization from posed multi-view inputs, but scaling them to large input view sets remains challenging. Transformer-based approaches that jointly process all input-view tokens incur rapidly increasing computation and memory as the number of views grows, while simple view subsampling discards potentially useful observations. We introduce Anchor-Integrated Multi-View Synthesis (AIMS), a scalable framework that decouples the number of available observations from the number of views processed by the global synthesis model. AIMS selects a fixed set of spatially distributed anchor views using farthest point sampling, groups nearby observations around each anchor, and uses a lightweight learnable integrator to fuse their information into enriched anchor representations. This allows additional observations to contribute to synthesis while keeping the downstream global view budget fixed. Evaluations on RealEstate10K and ScanNet demonstrate a favorable quality--efficiency trade-off against transformer-based and Gaussian-based baselines. AIMS achieves 29.41 dB and 17.73 dB PSNR on the two datasets, respectively, with rendering averaging 7.24 ms per view.