RSJEV: Discriminative Remote Sensing Scene Classification with Multimodal Large Language Models

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing remote sensing classification methods, which rely on predefined labels or inefficient autoregressive generation and struggle to balance accuracy with efficiency. We propose RSJEV, a framework that reformulates scene classification as a candidate-conditioned multimodal discriminative decision process by jointly modeling visual features, instructions, and category semantics. Central to this design is the OnePass Decider, which extracts multimodal decision states and directly estimates category probabilities, thereby eliminating autoregressive decoding while preserving vision-language interaction. Despite its compact 0.8B architecture, RSJEV outperforms mainstream methods across three benchmarks, substantially reducing inference costs and achieving a superior accuracy-efficiency trade-off.
📝 Abstract
Remote sensing scene classification is a fundamental task in Earth observation and geospatial analysis. Existing approaches mainly follow three paradigms: task-specific visual classification, vision-language similarity matching, and autoregressive multimodal generation. However, visual classifiers rely on predefined label spaces, CLIP-based methods perform recognition through static image-text alignment, and multimodal large language models (MLLMs) introduce unnecessary token-level generation for classification tasks with explicit candidate categories. To address these limitations, we propose RSJEV, a one-pass multimodal decision framework for remote sensing scene classification. Unlike conventional MLLMs that formulate classification as autoregressive text generation, RSJEV reformulates scene classification as a candidate-conditioned multimodal discriminative decision process, where visual representations, task instructions, and candidate category semantics are jointly modeled. Specifically, we introduce a OnePass Decider that extracts multimodal decision states and directly estimates category probabilities within the candidate category space, eliminating autoregressive decoding while preserving vision-language interactions. Extensive experiments on three widely used remote sensing scene classification benchmarks, including UC Merced, AID, and NWPU-RESISC45, demonstrate that RSJEV achieves superior classification performance compared with representative CNN-, Transformer-, Mamba-, CLIP-, and MLLM-based methods. Moreover, RSJEV significantly reduces inference costs and achieves a better accuracy-efficiency trade-off with only a compact 0.8B-parameter model. These results demonstrate the effectiveness of state-conditioned multimodal decision making for efficient remote sensing image understanding. The code will be available at https://github.com/Dongtcs/RSJEV.
Problem

Research questions and friction points this paper is trying to address.

Remote Sensing Scene Classification
Multimodal Large Language Models
Autoregressive Generation
Discriminative Decision
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Large Language Models
Remote Sensing Scene Classification
Discriminative Decision
OnePass Decider
Autoregressive-free
🔎 Similar Papers
No similar papers found.
D
Dongchen Si
School of Computer Science, Wuhan University, Wuhan 430072, China
Di Wang
Di Wang
School of Computer Science, Wuhan University
Remote SensingDeep LearningComputer VisionHyperspectral Image Clasification
M
Mingzhen Xu
School of Computer Science, Wuhan University, Wuhan 430072, China
J
Jing Zhang
School of Computer Science, Wuhan University, Wuhan 430072, China
Bo Du
Bo Du
Department of Management, Griffith Business School
Sustainable TransportTravel BehaviourUrban Data AnalyticsLogistics and Supply Chain
L
Liangpei Zhang
State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University, Wuhan 430079, China