🤖 AI Summary
This study addresses the challenge that structural discrepancies across state-level traffic accident databases hinder the reuse of predictive workflows, while rare fatal accidents remain poorly identified. To this end, it transfers a large language model (LLM)-based accident severity prediction pipeline to Tennessee data, employing textual prompts to handle heterogeneous attributes and fine-tuning Llama 3.1 8B via LoRA for classification. Furthermore, an evaluation framework integrating sampling strategies with class balancing is proposed. The results reveal that aggregated metrics can obscure performance deficiencies in rare categories; under random sampling, the weighted F1-score approaches 80%, yet the F1-score for fatal accidents remains below 27%. Introducing balanced sampling substantially improves the fatal-accident F1-score to 65.9%, achieving more equitable classification performance.
📝 Abstract
State crash databases differ in structure, coding, and injury-severity distributions, limiting direct reuse of predictive workflows across jurisdictions. This study adapts the SafeTraffic Copilot large language model (LLM) crash-severity workflow to a three-year Tennessee inventory of 624,392 crashes. Tennessee crash, roadway, vehicle, and person attributes were harmonized and converted into textual prompts while unavailable values were preserved rather than inferred. Llama 3.1 8B was fine-tuned using low-rank adaptation to classify five injury-severity categories. Random, county-based, and severity-balanced sampling strategies were evaluated using separate in-sample and unseen-test experiments; unseen-test experiments used 70/15/15 training, validation, and test splits. On their respective unseen test sets, random and county-based models achieved weighted F1-scores near 80%, but macro F1 remained below 47% and fatal-crash F1 below 27%, showing that strong aggregate performance can mask weak recognition of rare outcomes. Within its balanced evaluation population, the severity-balanced model achieved weighted and macro F1-scores of 57.7% and a fatalcrash F1-score of 65.9%, yielding more even class-level performance. Because each sampling strategy used a different test subset, crossstrategy differences are descriptive rather than controlled rankings. The results highlight the importance of interpreting LLM crashseverity performance together with sampling design, class balance, class-level metrics, and evaluation-population composition.