🤖 AI Summary
This study addresses the ray bundle ambiguity of visual tokens caused by projection discrepancies in feed-forward novel view synthesis with heterogeneous cameras. To this end, it proposes a unified ray-space framework that models perspective, fisheye, and panoramic cameras as calibrated samples within a shared ray space. Central to this approach are a local ray-graph representation and a projection-aware 2D RoPE, which jointly leverage token-centric relative positional encoding and ray-based angular coordinate reasoning to eliminate cross-sensor geometric inconsistencies. Evaluations on the ScanNet++ mixed-camera benchmark demonstrate that the proposed method significantly outperforms existing baselines while achieving zero-shot generalization to panoramic views.
📝 Abstract
Feed-forward novel view synthesis has recently shown promising results from sparse posed images, but most existing methods assume that context and target views share a fixed camera family. This homogeneous-camera assumption breaks in practical multi-sensor systems, where perspective, fisheye, and panoramic cameras may coexist and where the target projection may be unseen during training. We study feed-forward NVS across heterogeneous central cameras and identify a key ambiguity introduced by tokenization: a visual token aggregates a projection-dependent bundle of pixel rays, while existing camera encodings mainly expose absolute rays or token-center relations. To address this, we combine token-center relative Camera Positional Encodings and proposed local raymaps, a token-level representation that explicitly describes the intra-patch ray distribution summarized by each token. We further propose projection-aware 2D RoPE, which replaces raw image-grid coordinates with ray-induced angular coordinates so that relative positional reasoning is aligned across camera projections. Together, these components treat diverse cameras as calibrated samplings of a shared ray space rather than separate visual domains. On ScanNet++ with heterogeneous-camera system, our method improves over camera-conditioned baselines under mixed-camera evaluation and demonstrates zero-shot generalization to panoramic views.