🤖 AI Summary
This study addresses the challenge of analyzing complex, multidimensional safety risks in aviation systems, which traditional functional hazard analysis methods struggle to capture comprehensively. To overcome this limitation, the authors propose a novel, traceable, and structured approach for generating hypothetical hazard scenarios using large language models (LLMs). The method automatically constructs coherent hazard narratives from NASA Aviation Safety Reporting System (ASRS) reports and evaluates their plausibility through historical co-occurrence evidence. Innovatively integrating evolutionary abduction with a hybrid generation mechanism—combining zero-shot and few-shot prompting alongside optional fine-tuning—the framework employs evolutionary algorithms to optimize both structural validity and narrative consistency. Experimental results demonstrate that this hybrid strategy significantly enhances the realism, logical correctness, and diversity of generated scenarios, outperforming approaches based on single-generation paradigms.
📝 Abstract
Operational hazard analysis of aviation system operations must consider interactions among weather, ATC actions, airspace constraints, aircraft operations, and human factors - distinct from the functional hazard assessment applied at the aircraft-system level. We present an AI-assisted approach that generates candidate hazard scenarios from NASA's Aviation Safety Reporting System (ASRS). Given a target adverse outcome, it produces a structured hypothesis as categorical factors and a narrative scenario describing an operational event sequence consistent with the structure. Each scenario includes by a plausibility score from historical co-occurrence evidence and traceability to the most similar held-out ASRS reports. We then propose a hybrid variant, conditioning narrative generation on a structured hypothesis produced via evolutionary abduction, improving correctness and reducing variability. We evaluate multiple large language models, zero-shot versus few-shot prompting, and optional fine-tuning, measuring how prompting and model choice affect the validity and realism of the generated structures and narratives.