🤖 AI Summary
This study addresses the challenges of small object recognition and reliance on external tools for reasoning in ultra-high-resolution remote sensing visual question answering (VQA) by proposing an efficient framework based on privileged self-distillation. Methodologically, it designs context-preserving visual privileges and a correctness-aligned distillation mechanism to internalize zooming capabilities while avoiding information loss and teacher signal conflicts. Additionally, the GeoEvidence-6K dataset is constructed, integrating human feedback to guide skill refinement and employing online policy self-distillation for model optimization. Experimental results demonstrate that the proposed approach achieves state-of-the-art performance across multiple benchmarks with an average improvement of 4.0%. Notably, the lightweight 2B model surpasses most 8B counterparts while attaining the fastest inference speed.
📝 Abstract
Ultra-high-resolution (UHR) remote sensing visual question answering (VQA) requires models to resolve small visual evidence within extremely large images. Existing approaches typically rely on token pruning, visual search, or tool-augmented reasoning at inference time. We instead investigate whether the benefit of zoom-in visual privilege can be internalized into the model. We introduce RS-OPSD, a reliable privileged on-policy self-distillation (OPSD) framework for UHR remote sensing VQA. To provide high-quality privileged information with explicit question-relevant evidence, we construct GeoEvidence-6K, containing 6,750 VQA samples across seven task categories with evidence-region annotations, and develop Human Feedback-Guided Skill Refinement (HF-SR) for scalable annotation. To address context loss from tight crops and conflicting signals from imperfect teachers, RS-OPSD introduces Context-Preserving Visual Privilege (CPVP) and Correctness-Aligned Distillation (CAD). Without any additional visual search or tool calls at inference time, RS-OPSD achieves state-of-the-art (SOTA) performance on XLRS-Bench, MME-RealWorld-RS, and LRS-VQA, outperforming pervious SOTA models of comparable scale by an average of 4.0 percentage points. Moreover, our 2B variant, RS-OPD-Lite, surpasses most 8B-scale models while achieving the fastest measured inference speed. Our Code, GeoEvidence-6K, and the model weights for RS-OPSD and RS-OPD-Lite are publicly available.