RT-DETR-World: Transferring Rich LLM Semantics to Real-Time Open-Vocabulary Detection

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insufficient semantic absorption capacity of compact models in real-time open-vocabulary detection, where balancing generalization and efficiency remains challenging. We propose a lightweight DETR-based detection method incorporating Dual-path Description Alignment (DDA) and Relation-aware Negative Relaxation (RNR). By constructing the GroundingCapv2 dataset with LLM-generated descriptions as supervisory signals, rich semantic knowledge from large models is distilled into a lightweight detector during training while remaining fully decoupled at inference. Integrating MiniLM encoding with multi-granularity vision-language alignment, our approach achieves highly competitive zero-shot detection accuracy, striking an effective balance between precision and inference efficiency. The code will be made publicly available soon.
📝 Abstract
Open-vocabulary detection (OVD) recognizes categories unseen during training through textual category queries, yet achieving strong generalization with real-time efficiency remains challenging. Beyond vocabulary scaling, zero-shot generalization may benefit from reusable visual--semantic cues learned from seen data, including attributes, actions, states, and contextual relations. Existing real-time OVD methods primarily emphasize vocabulary coverage and efficient region/query--text matching; under strict efficiency constraints, compact detectors may struggle to absorb rich instance semantics and scene context. We propose RT-DETR-World, a compact DETR-style detector that transfers the rich semantics conveyed by descriptions during training while retaining lightweight query--text matching at inference. We construct GroundingCapv2 with three levels of supervision: category names for standard OVD, object descriptions conveying instance-level semantics, and image descriptions conveying object relations and scene context. These descriptions serve only as training-time semantic supervision. To help the compact detector absorb these semantics, we propose Dual-Path Description Alignment (DDA), combining a deployment-consistent MiniLM pathway with a training-only LLM teacher. MiniLM provides query--category supervision and object-description alignment, while offline teacher features supervise matched queries and global visual representations at the object and image levels, respectively. All teacher features are precomputed, and the teacher-side modules are removed after training. We further propose Relation-Aware Negative Relaxation (RNR), which uses teacher-derived semantic similarities to relax related negatives while preserving exact positives. Experiments demonstrate competitive zero-shot accuracy and a favorable accuracy--efficiency trade-off. The code will be released.
Problem

Research questions and friction points this paper is trying to address.

Open-vocabulary detection
Real-time efficiency
Zero-shot generalization
Instance semantics
Scene context
Innovation

Methods, ideas, or system contributions that make the work stand out.

Open-Vocabulary Detection
Real-Time Detection
Knowledge Distillation
Description Alignment
Negative Relaxation