GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that visual localization in complex scenes fails to meet the real-time and precision requirements of closed-loop robotic control. To this end, this work proposes a 4B-parameter foundation model that establishes precise localization as the perceptual cornerstone of physical intelligence. Methodologically, it constructs a unified framework integrating multimodal spatial pre-training, supervised fine-tuning, and GRPO-based reinforcement learning. The approach employs a shared vocabulary to generate quantized point and box coordinates while introducing dense localization and OCR data to enhance feature representations. Evaluated across 34 benchmarks, the proposed model achieves an average accuracy of 73.68%, establishing new state-of-the-art records. Furthermore, it significantly improves performance in downstream applications such as robotics and autonomous driving.
📝 Abstract
Precise grounding matters. It specifies which object is the target and where that object is, even in clutter and for tiny objects, and it has to be fast enough for closed-loop control. Yet vision-language-action (VLA) and world-action models (WAMs) take perception from general-purpose vision-language and video-generation backbones, which still fail in these settings. We introduce GroundingPI, a 4B grounding foundation model that generates points and boxes as quantized coordinates in a shared vocabulary. Training combines multimodal and spatial pretraining, supervised fine-tuning, and reinforcement learning with GRPO, using supervision from public datasets and dedicated data engines. Against 44 baselines across 34 grounding benchmarks spanning 11 perceptual capabilities, GroundingPI establishes a new state of the art, averaging 73.68%, above the larger GPT-6 Astra (71.54%). As a downstream visual backbone, GroundingPI improves performance on robotic manipulation and autonomous driving. On RoboTwin 2.0, it outperforms every mainstream backbone we evaluate in all four out-of-distribution settings, by up to 24.8% relative to the strongest backbone. On RoboCasa-GR1, GroundingPI trained with 50% of the demonstrations outperforms those baselines trained with 75%. On nuScenes, used as the visual backbone, GroundingPI attains an average open-loop L2 error of 0.296 m. We systematically analyze GroundingPI's pretraining in scale and data composition. Downstream autonomous driving and robotic manipulation improve as the pretraining is scaled. Analyzing the data recipe across these 11 perceptual capabilities shows dense grounding's substantial benefits for both, and OCR's potential as a catalyst for perceptual learning. These results support grounding as a perceptual foundation, and dedicated perceptual pretraining as a promising direction for foundation models of physical intelligence.
Problem

Research questions and friction points this paper is trying to address.

visual grounding
physical intelligence
vision-language-action models
robotic manipulation
autonomous driving
Innovation

Methods, ideas, or system contributions that make the work stand out.

Grounding Foundation Model
Physical Intelligence
Visual Primitives
Reinforcement Learning with GRPO
Vision-Language-Action
💼 Related Jobs
No related jobs found.