A Unified Framework and Dataset for Oriented Object Visual Grounding in Remote Sensing

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
该研究解决了遥感图像中任意方向物体的视觉定位问题,通过引入O²-VG模型家族及DIOR-R-RSVG数据集,采用跨模态变换器和自回归视觉-语言模型方法。
📝 Abstract
Visual grounding in remote sensing images aims to locate objects described by referring expressions. Most existing methods predict horizontal bounding boxes, which are often inaccurate for objects with arbitrary orientations. To address this limitation, we introduce O$^2$-VG, a family of models for oriented object visual grounding with three complementary designs. Specifically, O$^2$-VG-Trans is a cross-modality transformer for oriented object visual grounding. It establishes a strong discriminative foundation for the model family. Building upon it, O$^2$-VG-Uni predicts universal oriented proposals for possible foreground objects without specific text prompts. It also supports object retrieval through cached proposal embeddings. Using these universal oriented proposals as input prompts, O$^2$-VG-VLM is an autoregressive vision-language model. It generates oriented box token blocks in parallel through multi-token prediction. In addition, we construct DIOR-R-RSVG, a dataset for oriented object visual grounding in remote sensing images. It provides image, expression, and oriented box triplets for training and evaluation. Together, the O$^2$-VG family provides a flexible framework that spans discriminative transformers and generative vision-language models. It achieves superior performance across multiple benchmarks. Code is available at https://github.com/wokaikaixinxin/ai4rs.
Problem

Research questions and friction points this paper is trying to address.

Visual Grounding
Remote Sensing
Oriented Objects
Innovation

Methods, ideas, or system contributions that make the work stand out.

Oriented Object Visual Grounding
Cross-modality Transformer
Universal Oriented Proposals
Autoregressive Vision-Language Model
DIOR-R-RSVG Dataset
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Zeyu Ding
Zeyu Ding
Binghamton University
privacysecuritymachine learningfairnessformal verification
Y
Yong Zhou
School of Computer Science and Technology/School of Artificial Intelligence, the Mine Digitization Engineering Research Center of the Ministry of Education, and Jiangsu Provincial Industrial Technology Engineering Center for Intelligent Sensing and Emergency IoT in Underground Space, China University of Mining and Technology, Xuzhou 221116, China
Jiaqi Zhao
Jiaqi Zhao
Xidian University
privacy-preserving machine learning
W
Wen-Liang Du
School of Computer Science and Technology/School of Artificial Intelligence, the Mine Digitization Engineering Research Center of the Ministry of Education, and Jiangsu Provincial Industrial Technology Engineering Center for Intelligent Sensing and Emergency IoT in Underground Space, China University of Mining and Technology, Xuzhou 221116, China
X
Xixi Li
School of Computer Science and Technology/School of Artificial Intelligence, the Mine Digitization Engineering Research Center of the Ministry of Education, and Jiangsu Provincial Industrial Technology Engineering Center for Intelligent Sensing and Emergency IoT in Underground Space, China University of Mining and Technology, Xuzhou 221116, China
H
Hancheng Zhu
School of Computer Science and Technology/School of Artificial Intelligence, the Mine Digitization Engineering Research Center of the Ministry of Education, and Jiangsu Provincial Industrial Technology Engineering Center for Intelligent Sensing and Emergency IoT in Underground Space, China University of Mining and Technology, Xuzhou 221116, China
Rui Yao
Rui Yao
China University of Mining and Technology
Computer VisionMachine Learning
Abdulmotaleb El Saddik
Abdulmotaleb El Saddik
MCRLab, University of Ottawa
Immersive MediaDigital TwinsHuman Centered AIMultimedia CommunicationMetaverse