Traceable LLM-Generated Hazard Scenarios for Operational Safety Analysis of Aviation Systems Using ASRS Reports

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of analyzing complex, multidimensional safety risks in aviation systems, which traditional functional hazard analysis methods struggle to capture comprehensively. To overcome this limitation, the authors propose a novel, traceable, and structured approach for generating hypothetical hazard scenarios using large language models (LLMs). The method automatically constructs coherent hazard narratives from NASA Aviation Safety Reporting System (ASRS) reports and evaluates their plausibility through historical co-occurrence evidence. Innovatively integrating evolutionary abduction with a hybrid generation mechanism—combining zero-shot and few-shot prompting alongside optional fine-tuning—the framework employs evolutionary algorithms to optimize both structural validity and narrative consistency. Experimental results demonstrate that this hybrid strategy significantly enhances the realism, logical correctness, and diversity of generated scenarios, outperforming approaches based on single-generation paradigms.
📝 Abstract
Operational hazard analysis of aviation system operations must consider interactions among weather, ATC actions, airspace constraints, aircraft operations, and human factors - distinct from the functional hazard assessment applied at the aircraft-system level. We present an AI-assisted approach that generates candidate hazard scenarios from NASA's Aviation Safety Reporting System (ASRS). Given a target adverse outcome, it produces a structured hypothesis as categorical factors and a narrative scenario describing an operational event sequence consistent with the structure. Each scenario includes by a plausibility score from historical co-occurrence evidence and traceability to the most similar held-out ASRS reports. We then propose a hybrid variant, conditioning narrative generation on a structured hypothesis produced via evolutionary abduction, improving correctness and reducing variability. We evaluate multiple large language models, zero-shot versus few-shot prompting, and optional fine-tuning, measuring how prompting and model choice affect the validity and realism of the generated structures and narratives.
Problem

Research questions and friction points this paper is trying to address.

hazard scenario generation
aviation safety
operational hazard analysis
traceability
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

traceable hazard scenario generation
large language models
evolutionary abduction
operational safety analysis
ASRS