Adapting a Large Language Model Crash-Severity Pipeline to Tennessee: Performance Across Sampling Strategies

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that structural discrepancies across state-level traffic accident databases hinder the reuse of predictive workflows, while rare fatal accidents remain poorly identified. To this end, it transfers a large language model (LLM)-based accident severity prediction pipeline to Tennessee data, employing textual prompts to handle heterogeneous attributes and fine-tuning Llama 3.1 8B via LoRA for classification. Furthermore, an evaluation framework integrating sampling strategies with class balancing is proposed. The results reveal that aggregated metrics can obscure performance deficiencies in rare categories; under random sampling, the weighted F1-score approaches 80%, yet the F1-score for fatal accidents remains below 27%. Introducing balanced sampling substantially improves the fatal-accident F1-score to 65.9%, achieving more equitable classification performance.
📝 Abstract
State crash databases differ in structure, coding, and injury-severity distributions, limiting direct reuse of predictive workflows across jurisdictions. This study adapts the SafeTraffic Copilot large language model (LLM) crash-severity workflow to a three-year Tennessee inventory of 624,392 crashes. Tennessee crash, roadway, vehicle, and person attributes were harmonized and converted into textual prompts while unavailable values were preserved rather than inferred. Llama 3.1 8B was fine-tuned using low-rank adaptation to classify five injury-severity categories. Random, county-based, and severity-balanced sampling strategies were evaluated using separate in-sample and unseen-test experiments; unseen-test experiments used 70/15/15 training, validation, and test splits. On their respective unseen test sets, random and county-based models achieved weighted F1-scores near 80%, but macro F1 remained below 47% and fatal-crash F1 below 27%, showing that strong aggregate performance can mask weak recognition of rare outcomes. Within its balanced evaluation population, the severity-balanced model achieved weighted and macro F1-scores of 57.7% and a fatalcrash F1-score of 65.9%, yielding more even class-level performance. Because each sampling strategy used a different test subset, crossstrategy differences are descriptive rather than controlled rankings. The results highlight the importance of interpreting LLM crashseverity performance together with sampling design, class balance, class-level metrics, and evaluation-population composition.
Problem

Research questions and friction points this paper is trying to address.

crash severity prediction
large language models
sampling strategies
class imbalance
cross-jurisdictional adaptation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Model
Low-Rank Adaptation (LoRA)
Crash Severity Classification
Sampling Strategies
Prompt Engineering
🔎 Similar Papers
No similar papers found.
A
Abhilasha Saroj
Oak Ridge National Laboratory
P
Pranav Govindu
University of Tennessee at Knoxville
Bharat Sharma
Bharat Sharma
Oak Ridge National Laboratory
FATESCarbon Cycle and extremes Climate ExtremesBiogeochemistryEarth System ModelingSpace
U
Usman Ahmed
University of Tennessee - Oak Ridge Innovation Institute