🤖 AI Summary
This study addresses the reliability limitations of symmetric fusion in disaster assessment caused by cross-view conflicts by proposing a visibility-conditioned reliability gating mechanism. This approach reframes the fusion decision from whether to fuse into when and why to trust specific views. It leverages single-view models to identify conflicting samples across views, dynamically allocates view-specific weights via linear gating, and combines confidence calibration with field-of-view control experiments to reveal how image cropping affects fusion gains. Evaluated on a wildfire dataset, the proposed method significantly outperforms baselines (p<0.001). Furthermore, conflict density enables unsupervised prediction of tile-level damage (r=0.615), establishing a reliable new paradigm for multi-view disaster assessment.
📝 Abstract
After a disaster, building damage is assessed from overhead tiles and ground-level photographs, and most methods fuse the two views symmetrically, trusting both equally for every building. This paper focuses on the samples where that assumption fails: the conflict cases, on which two independently trained single-view models disagree. We mine such cases from three paired collections (inspection photographs from the 2025 Eaton wildfire and street-view panoramas from Hurricanes Ian and Milton, each matched to very-high-resolution overhead tiles), where they make up 10-33% of the data. On these samples an oracle that simply trusts the correct view beats every fusion method we tested by 0.37-0.41 accuracy, and the gap survives longer training, calibration, and backbone changes. We recover part of it with a visibility-conditioned reliability gate: a linear model that decides which view to trust from building-visibility features, calibrated per-view confidences, and the disagreement itself. On the wildfire data the gate is the only method that significantly beats calibrated probability averaging (+0.051 on conflicts, p=0.0001) and end-to-end fusion (+0.072, p<10^-4); on the panoramic datasets it matches them. A controlled field-of-view experiment explains why: cropping panoramas toward the building doubles the benefit of fusion, whereas random crops of the same size do not. Finally, the spatial density of conflicts predicts tile-level damage without labels (Spearman r=0.615, p=0.001). Mining conflicts turns"does fusion help?"into"which view should be trusted, where, and why?".