🤖 AI Summary
This work addresses the serialization latency and causal ordering constraints inherent in autoregressive visual grounding models by proposing a 4B-parameter foundation model. Methodologically, it introduces bidirectional diffusion into visual grounding for the first time, employing an autoregressive-to-diffusion transition combined with Group Relative Policy Optimization (GRPO) reinforcement learning. Block-level denoising is adopted to balance parallel decoding with precise localization, further integrated with entropy-guided decoding, multi-objective joint training, and progressive inference optimization. The proposed model establishes new state-of-the-art performance across 30 benchmarks, achieving an average accuracy of 72.42% while accelerating inference speed by 4.51×.
📝 Abstract
Autoregressive (AR) grounding models serialize spatial predictions, introducing sequential latency and imposing a causal order on output tokens. We view grounding as visual evidence extraction: objects, locations, and spatial relations are jointly constrained by the image and query, yet their dependencies do not imply an intrinsic left-to-right generation order. This distinction makes bidirectional diffusion a natural fit, allowing spatial hypotheses to emerge in parallel and be jointly refined through iterative denoising. We introduce GroundAnything, a 4B-parameter grounding foundation model that reconciles fast parallel decoding with precise localization through blockwise denoising. Training combines grounding pretraining from public datasets and dedicated data engines, direct AR-to-diffusion conversion with joint AR and diffusion objectives, supervised fine-tuning, and GRPO-based reinforcement post-training. Across 30 grounding benchmarks, our autoregressive variant, GroundAnything-VLM, establishes a new overall state of the art among similarly sized models at 72.42%, remaining competitive with GPT-6 Astra (71.35%). With entropy-guided decoding, GroundAnything also surpasses the prior state of the art at this scale, averaging 61.75% versus 53.32% for the fast MTP-based LocateAnything model. We further explore decoding strategies, showing that an optional self-speculative mode achieves a $4.51\times$ speedup over the AR counterpart with a 0.74 percentage-point drop in COCO F1mIoU. Infrastructure experiments show that progressive inference optimizations translate parallel decoding into practical speedups. These support efficient visual grounding in latency-sensitive real-world systems.