RS-OPSD: Reliable Privileged On-Policy-Self-Distillation for Ultra-High-Resolution Remote Sensing VQA

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of small object recognition and reliance on external tools for reasoning in ultra-high-resolution remote sensing visual question answering (VQA) by proposing an efficient framework based on privileged self-distillation. Methodologically, it designs context-preserving visual privileges and a correctness-aligned distillation mechanism to internalize zooming capabilities while avoiding information loss and teacher signal conflicts. Additionally, the GeoEvidence-6K dataset is constructed, integrating human feedback to guide skill refinement and employing online policy self-distillation for model optimization. Experimental results demonstrate that the proposed approach achieves state-of-the-art performance across multiple benchmarks with an average improvement of 4.0%. Notably, the lightweight 2B model surpasses most 8B counterparts while attaining the fastest inference speed.
📝 Abstract
Ultra-high-resolution (UHR) remote sensing visual question answering (VQA) requires models to resolve small visual evidence within extremely large images. Existing approaches typically rely on token pruning, visual search, or tool-augmented reasoning at inference time. We instead investigate whether the benefit of zoom-in visual privilege can be internalized into the model. We introduce RS-OPSD, a reliable privileged on-policy self-distillation (OPSD) framework for UHR remote sensing VQA. To provide high-quality privileged information with explicit question-relevant evidence, we construct GeoEvidence-6K, containing 6,750 VQA samples across seven task categories with evidence-region annotations, and develop Human Feedback-Guided Skill Refinement (HF-SR) for scalable annotation. To address context loss from tight crops and conflicting signals from imperfect teachers, RS-OPSD introduces Context-Preserving Visual Privilege (CPVP) and Correctness-Aligned Distillation (CAD). Without any additional visual search or tool calls at inference time, RS-OPSD achieves state-of-the-art (SOTA) performance on XLRS-Bench, MME-RealWorld-RS, and LRS-VQA, outperforming pervious SOTA models of comparable scale by an average of 4.0 percentage points. Moreover, our 2B variant, RS-OPD-Lite, surpasses most 8B-scale models while achieving the fastest measured inference speed. Our Code, GeoEvidence-6K, and the model weights for RS-OPSD and RS-OPD-Lite are publicly available.
Problem

Research questions and friction points this paper is trying to address.

Ultra-High-Resolution Remote Sensing VQA
Visual Privilege Internalization
Context Loss
Self-Distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Self-Distillation
Ultra-High-Resolution VQA
Visual Privilege
Context-Preserving
Correctness-Aligned Distillation
💼 Related Jobs
No related jobs found.
C
Chengjie Jiang
Tsinghua University
Y
Yunqi Zhou
Zhejiang University
J
Jiafeng Yan
Central University of Finance and Economics
S
Sihang Zhao
East China Normal University; Key Laboratory of Geographic Information Science
Chun Yuan
Chun Yuan
Graduate School at Shenzhen, Tsinghua University
Computer visionmultimedia access control
Jing Li
Jing Li
Institute of Automation, Chinese Academy of Sciences (CASIA), Beijing, 100190, China.
segmentationdeep learningweakly supervised learning