🤖 AI Summary
This study addresses the dual challenges of scarce real-world annotated videos and prohibitive high-fidelity simulation costs in wildfire monitoring. To overcome these limitations, this work proposes a simulation-based multi-agent vision-language model (VLM) framework that transforms 2D wildfire simulations into labeled videos via Blender mapping and controllable generation, incorporating an automated simulation-to-agent conversion mechanism. By integrating training-free multi-agent collaboration with retrieval-augmented generation, the system enables video memory reasoning and automated structured report generation. This approach effectively mitigates the data scarcity bottleneck, achieving 51.5% four-label accuracy on video memory tasks and 77.3% accuracy across six report fields for the complete pipeline, substantially outperforming existing baseline methods.
📝 Abstract
Effective wildfire monitoring requires relating visual evidence to physical fire dynamics, yet real videos with synchronized physical annotations are scarce and high-fidelity 3D simulation is costly. We present a simulation-grounded vision-language model (VLM) framework that automatically converts 2D wildfire simulations into labeled video episodes. A fixed Blender mapping produces low-detail 3D proxies aligned with simulator terrain, fuel layout, fire activity, and wind cues; controllable video generation supplies richer appearance. The proxies are intermediate representations rather than finely rendered final scenes. Generated videos and simulator labels form reusable multimodal memory for a training-free multi-agent VLM system that retrieves reference episodes, reconciles visual and memory-based predictions, and produces structured wildfire reports. On held-out generated episodes, video memory achieves 51.5% exact four-tag accuracy, compared with 22.6% for direct VLM querying and 16-17% for text-only memory; the complete system achieves 77.3% accuracy on six simulator-derived report fields. Component ablations, cross-generator tests, and three real-UAV evaluations assess retrieval, reporting, generator changes, and observable monitoring tasks. The framework connects automatic simulation-to-proxy conversion with memory-based VLM reasoning under scarce real-world physical annotations.