GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the serialization latency and causal ordering constraints inherent in autoregressive visual grounding models by proposing a 4B-parameter foundation model. Methodologically, it introduces bidirectional diffusion into visual grounding for the first time, employing an autoregressive-to-diffusion transition combined with Group Relative Policy Optimization (GRPO) reinforcement learning. Block-level denoising is adopted to balance parallel decoding with precise localization, further integrated with entropy-guided decoding, multi-objective joint training, and progressive inference optimization. The proposed model establishes new state-of-the-art performance across 30 benchmarks, achieving an average accuracy of 72.42% while accelerating inference speed by 4.51×.
📝 Abstract
Autoregressive (AR) grounding models serialize spatial predictions, introducing sequential latency and imposing a causal order on output tokens. We view grounding as visual evidence extraction: objects, locations, and spatial relations are jointly constrained by the image and query, yet their dependencies do not imply an intrinsic left-to-right generation order. This distinction makes bidirectional diffusion a natural fit, allowing spatial hypotheses to emerge in parallel and be jointly refined through iterative denoising. We introduce GroundAnything, a 4B-parameter grounding foundation model that reconciles fast parallel decoding with precise localization through blockwise denoising. Training combines grounding pretraining from public datasets and dedicated data engines, direct AR-to-diffusion conversion with joint AR and diffusion objectives, supervised fine-tuning, and GRPO-based reinforcement post-training. Across 30 grounding benchmarks, our autoregressive variant, GroundAnything-VLM, establishes a new overall state of the art among similarly sized models at 72.42%, remaining competitive with GPT-6 Astra (71.35%). With entropy-guided decoding, GroundAnything also surpasses the prior state of the art at this scale, averaging 61.75% versus 53.32% for the fast MTP-based LocateAnything model. We further explore decoding strategies, showing that an optional self-speculative mode achieves a $4.51\times$ speedup over the AR counterpart with a 0.74 percentage-point drop in COCO F1mIoU. Infrastructure experiments show that progressive inference optimizations translate parallel decoding into practical speedups. These support efficient visual grounding in latency-sensitive real-world systems.
Problem

Research questions and friction points this paper is trying to address.

Visual Grounding
Autoregressive Models
Parallel Decoding
Sequential Latency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Parallel Decoding
Visual Grounding
Bidirectional Diffusion
Blockwise Denoising
Self-Speculative Decoding