🤖 AI Summary
This study addresses the critical risk that over-optimistic evaluations of large language models (LLMs) in wildfire management may lead to severe casualties and property damage, necessitating rigorous reliability validation under real-world conditions. We conduct zero-shot evaluations of 35 LLMs across five wildfire-related tasks, introducing a novel "unassisted–information-augmented" paired experimental paradigm that enhances reasoning through domain knowledge such as read-only SQL tools and reference frames. Our findings reveal key risks, including that simple heuristics frequently outperform model predictions and that public datasets are susceptible to label leakage. This work represents the first integration of large-scale model benchmarking, comprehensive task coverage, and dual-setting evaluation in this domain. All code and results have been publicly released to facilitate future research.
📝 Abstract
Large language models (LLMs) are entering wildfire management, where overstated evaluations can cost property and lives. How do they perform on wildfire tasks, with and without grounding? Bare means a model receives the task input alone. Grounded means it also receives one task-specific addition: for smoke detection, a smoke-free reference frame from the same camera. AI4Fire runs six core models bare and grounded on five wildfire tasks, zero-shot; a sweep adds 29 more. Our literature search on fire tasks found 138 works; none combines this roster, task coverage, and paired bare and grounded runs. We report three findings. (1) Grounding helped most where the addition carried the answer: a read-only SQL tool lifted every core model's database accuracy from at most 16 to at least 88 percent. (2) Simple rules were hard to beat: no core model outperformed repeating today's staffing count, and two open-weight models mostly copied the median of similar earlier fire-days, a worse forecast. (3) Public releases carry hazards: a fire-danger column separates the holdout perfectly, and 67 aerial fire frames carry smoldering or fire-free labels read from a clipped thermal maximum. We release prompts, responses, scores, code, and the survey record.