GeoArbiter: Verifiability-Guided Grounding for Remote-Sensing Multimodal LLMs

📅 2026-08-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the susceptibility of multimodal large language models (MLLMs) in remote sensing to generate hallucinations inconsistent with input imagery and to be misled by conflicting geospatial retrieval results. To mitigate these issues, the authors propose GeoArbiter, a training-free inference pipeline that introduces, for the first time, a cross-modal verifiability-based content filtering mechanism. GeoArbiter dynamically assesses retrieved geographic knowledge through coordinate-keyed retrieval and content-level filtering, selectively injecting only those facts that are reliable yet unverifiable from the image itself, thereby avoiding attribute confusion and response bias. Evaluated on three open-source remote sensing MLLMs, GeoArbiter preserves 84.69–87.15% of retrieval gains while reducing hallucinations by 9.58–26.34% and significantly enhancing robustness against conflicting external information.
📝 Abstract
Remote-sensing multimodal large language models (MLLMs) often assert facts that imagery cannot establish, such as a facility's identity or function. Coordinate-keyed geographic retrieval can supply this missing knowledge, improving fMoW land-use accuracy by 12.06--17.19 points across three open MLLMs. However, retrieved records can also contradict visible evidence, and we find that models frequently follow the records even when the image is decisive. We argue that source trust should therefore depend on \emph{cross-modal verifiability}: geographic records are most useful for attributes the image cannot verify and most dangerous when they dispute visually verifiable attributes. We introduce GeoArbiter, a training-free pipeline that operationalizes this principle by injecting only image-unverifiable geographic facts. Unlike arbitration prompts, which leak across attribute types and bias yes/no responses, content-level filtering preserves 84.69--87.15\% of the full-retrieval accuracy gain, reduces claim-level hallucination by 9.58--26.34\% under a source-blinded judge, and improves robustness to conflicting records across all three models. These results identify verifiability-guided content selection as a simple, effective mechanism for grounding remote-sensing MLLMs in fallible geographic knowledge.
Problem

Research questions and friction points this paper is trying to address.

remote-sensing MLLMs
geographic retrieval
cross-modal verifiability
hallucination
grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-modal verifiability
GeoArbiter
remote-sensing MLLMs
hallucination reduction
geographic knowledge grounding
🔎 Similar Papers
No similar papers found.