Deep Prior Learning for Embodied Perception

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of noisy pose handling, prior dilution, and physical scale recovery in embodied perception by proposing the VPGGT framework. Methodologically, it introduces a parameter-free prior residual connection to mitigate prior dilution in deep Transformer layers, and designs a pose-depth-conditioned metric global attention mechanism that predicts a shared scaling factor for geometric scale recovery. Furthermore, sensor-driven pose corruption and feature-level conditioning strategies are incorporated to enhance robustness. Experiments across four datasets demonstrate that the proposed framework significantly improves translation accuracy and joint pose AUC, validating the effectiveness of explicit prior access.
📝 Abstract
Embodied systems need geometric perception that exploits available observations beyond images alone. Recent feed-forward 3D models incorporate geometric priors, including camera poses, intrinsics, and depth. However, handling noisy poses, preserving accurate priors, and recovering physical scale require more than simply accepting these inputs. We introduce \emph{Vision-Prior Geometry Grounded Transformer} (VPGGT), a VGGT-based framework that extends OmniVGGT for prior-aware embodied perception. We formulate sensor-motivated pose corruptions from ground-truth trajectories for training and introduce a parameter-free \emph{prior residual connection} (PRC) to mitigate \emph{prior dilution}, where predictions are less accurate than their supplied pose priors. Our noise formulation targets camera poses; supplied intrinsics and depth receive no additional corruption. We further introduce \emph{Metric Global Attention}, which conditions a global scale token on available pose and depth scales and predicts a shared metric scaling factor for the geometric outputs. Experiments across four datasets show that \emph{PRC} improves translation-direction accuracy and joint pose AUC over a matched training baseline when camera priors are provided for all views, under both exact and corrupted poses. These results support explicit prior access during refinement as a useful addition to feature-level conditioning.
Problem

Research questions and friction points this paper is trying to address.

embodied perception
geometric priors
noisy poses
prior dilution
physical scale recovery
Innovation

Methods, ideas, or system contributions that make the work stand out.

Embodied Perception
Prior Residual Connection
Metric Global Attention
Geometry Grounded Transformer
Prior Dilution
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.