OmniAct3D: Leveraging Foundation Geometry and Evidence-Grounded Reasoning for Panoramic 3D Detection

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant geometric distortions inherent in equirectangular projection (ERP) for panoramic 360° scenes, which hinder the direct transfer of perspective-domain pre-trained vision foundation models to 3D detection tasks. To overcome this, we propose OmniAct3D, a novel framework that introduces the first ERP ray-geometry adapter to eliminate spherical projection discrepancies. Furthermore, it designs a visual-action evidential reasoning chain for structured geometric action generation and integrates multi-resolution feature re-encoding with an appearance-guided expert module to achieve efficient cross-perception feature alignment. Experiments on the Spheriverse and PanoMMOcc datasets demonstrate improvements of 2.96 and 24.87 points in NDS and mAP, respectively. Notably, the reasoning chain retains 95%–98% of baseline performance, enabling reusable object-level panoramic 3D reasoning.
📝 Abstract
Accurate 3D detection is essential for mobile embodied agents, while Vision Foundation Models (VFMs) offer transferable visual and geometric priors. Yet existing VFM-based 3D detectors rely on narrow-view monocular images or discrete perspective views, limiting coherent surround perception; equirectangular projection (ERP) instead encodes a continuous 360 scene in a single image. Direct transfer remains difficult because ERP organizes geometry and visual information differently, making object-relevant cues hard to model, localize, and preserve. We propose OmniAct3D, a framework that adapts perspective-trained VFM detectors to ERP while preserving transferable VFM priors. To resolve geometric mismatch, the ERP-Ray Geometry Adapter (ERGA-Ray) models spherical viewing rays and periodic spatial structure. To localize evidence in scene-wide context, the Visual-Action Reasoning Chain (VARC) grounds each hypothesis in relevant panoramic evidence and converts it into a structured geometric action. To recover local cues lost under fixed token budgets, the Appearance-Guided Heading Expert (AGHE) re-encodes object regions at higher resolution for heading estimation. Experiments show that OmniAct3D improves over the previous best 3D detector by 2.96 NDS points on Spheriverse and over the unadapted VFM baseline by 24.87 mAP points on PanoMMOcc. With target-specific geometry adaptation, VARC retains 95--98% of the same-configuration mAP, indicating reusable object-level 3D reasoning across sensing configurations. The source code will be made publicly available at https://github.com/FeiT-FeiTeng/OmniAct3D.
Problem

Research questions and friction points this paper is trying to address.

3D detection
Vision Foundation Models
equirectangular projection
panoramic perception
domain adaptation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Panoramic 3D Detection
Vision Foundation Models
Equirectangular Projection
Geometry Adapter
Visual-Action Reasoning Chain
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.