π€ AI Summary
This study addresses the feature mismatch between pre-trained 3D encoders designed for global scenes and the local observations encountered by embodied agents, proposing a pioneering label-free adaptation paradigm. The method freezes a global 3D encoder and trains only a lightweight adapter module using paired geometric information. Through point-feature alignment and relational distillation, it maps local-view features into the global semantic space, enabling global guidance during training while directly processing local observations at inference. Experiments demonstrate that this approach outperforms supervised parameter-efficient fine-tuning on the Sonata and Concerto benchmarks. Furthermore, its zero-shot cross-dataset transfer performance significantly surpasses that of fully fine-tuned models.
π Abstract
Pretrained 3D encoders are typically developed on globally reconstructed scenes expressed in a consistent world coordinate frame, whereas embodied systems must reason from partial, viewpoint-dependent observations in camera coordinates. We show that this shift from globally learned 3D feature spaces to realistic partial observations exposes a severe representation mismatch, which we find consistently across representative state-of-the-art encoders, including Sonata and Concerto. A frozen Sonata encoder with a global linear probe achieves 72.47 mIoU on full ScanNet scenes, but 2.57 mIoU on single-frame camera-coordinate inputs. Training-free gravity alignment recovers performance to 41.64 mIoU, showing that coordinate-frame mismatch is a dominant source of degradation but cannot be fully resolved through canonicalization alone. We introduce PAGER, a label-free adaptation method that aligns partial-view features with a frozen global 3D semantic space using only paired partial/global geometry. It learns lightweight adaptation modules while keeping the pretrained encoder and global segmentation probe frozen. Matched-point feature alignment anchors partial features to their global counterparts, while relational supervision preserves their similarity structure with respect to the global representation. Global geometry provides supervision only during training. Inference operates directly on the partial observation. Without partial-view labels, PAGER outperforms label-supervised PEFT on both Sonata and Concerto, and in zero-shot ScanNet$\rightarrow$ScanNet++ transfer surpasses fully fine-tuned Sonata ($53.93$ vs.\ $48.09$ mIoU), suggesting that preserving the frozen global representation can improve cross-dataset transfer.