🤖 AI Summary
This study addresses the trust crisis precipitated by AI-generated content in disaster-related social media, where the reliability of existing text detectors remains questionable. To investigate this, we propose four semantic matching-based data construction paradigms to build a disaster-domain text dataset. We systematically evaluate detection models and in-domain calibration methods through a comprehensive framework integrating LLM-based judgment, frozen-encoder linear probing, paired sensitivity analysis, and multi-model cross-validation. Our findings reveal that superficial feature biases fundamentally undermine detection efficacy, demonstrating that general-purpose detectors perform near chance level (AUROC ~0.5) in this domain. Consequently, we establish that single-modality text detection is unreliable for identifying AI-generated disaster content, necessitating the integration of multimodal evidence to implement effective trust gating mechanisms.
📝 Abstract
Disaster social sensing converts public social-media posts into evidence for situational awareness and humanitarian needs, but generative artificial intelligence (AI) can produce plausible messages that resemble eyewitness reports. This study investigates whether text-based AI detectors can reliably distinguish human-authored from AI-generated disaster posts. We construct a dataset of 12,000 texts organised into 3,000 matched semantic units from nine disasters: original human posts (H0), minimally LLM-proofread human posts (H1), factual AI-generated posts based on the same verified facts (A0), and affectively framed versions of those AI posts (A1). A separate 6,000-text corpus from 42 events supports model selection and threshold calibration. We evaluate OSM-Det, Fast-DetectGPT, Binoculars, and direct large language model (LLM) judges across five model families, then test disaster-domain calibration, a frozen-encoder linear readout, paired transformation sensitivity, and dataset artifact controls. Across fourteen frozen cross-family configurations, AUROC is 0.402-0.517 and the best prospective recall at a calibration-derived low-false-positive operating point is 3.6%; OSM-Det reaches AUROC 0.521 and 10.4% recall at a realised 6.7% false-positive rate. A disaster-trained linear head reaches AUROC 0.817, but a seven-feature surface classifier reaches 0.784 on the H0-versus-A0 contrast, and neutralising identified surface asymmetries reduces the head from 0.733 to 0.594. The head also separates A0 from A1 even though provenance is unchanged. The results show that text-based detection is not reliable enough to serve as an operational trust gate; multimodal claims, accountable sources, and other contextual evidence should be rested on to safeguard trust in disaster social sensing.