LIFT: Interpretable truck driving risk prediction with literature-informed fine-tuned LLMs

📅 2025-10-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the limited interpretability of models in truck driving risk prediction, this paper proposes LIFT—a literature-guided fine-tuning framework that systematically integrates an automatically constructed domain-specific knowledge base (covering traffic and safety literature) into large language model (LLM) training. LIFT unifies risk prediction and attribution analysis via fine-grained parameter tuning, variable importance ranking, and PERMANOVA-based statistical validation. Evaluated on real-world driving data, LIFT achieves a 26.7% improvement in recall and a 10.1% gain in F1-score over baseline models. Its explanations align with expert domain knowledge, exhibit robustness to data sampling perturbations, and successfully identify statistically validated high-risk variable combinations. This work represents the first systematic incorporation of structured scientific literature into LLM-driven risk prediction—thereby jointly advancing predictive accuracy, model interpretability, and empirical verifiability.

Technology Category

Natural Language Processing: Safety and RobustnessMachine Learning: Calibration & Uncertainty QuantificationReasoning under Uncertainty: Other Foundations of Reasoning under Uncertainty

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
This study proposes an interpretable prediction framework with literature-informed fine-tuned (LIFT) LLMs for truck driving risk prediction. The framework integrates an LLM-driven Inference Core that predicts and explains truck driving risk, a Literature Processing Pipeline that filters and summarizes domain-specific literature into a literature knowledge base, and a Result Evaluator that evaluates the prediction performance as well as the interpretability of the LIFT LLM. After fine-tuning on a real-world truck driving risk dataset, the LIFT LLM achieved accurate risk prediction, outperforming benchmark models by 26.7% in recall and 10.1% in F1-score. Furthermore, guided by the literature knowledge base automatically constructed from 299 domain papers, the LIFT LLM produced variable importance ranking consistent with that derived from the benchmark model, while demonstrating robustness in interpretation results to various data sampling conditions. The LIFT LLM also identified potential risky scenarios by detecting key combination of variables in truck driving risk, which were verified by PERMANOVA tests. Finally, we demonstrated the contribution of the literature knowledge base and the fine-tuning process in the interpretability of the LIFT LLM, and discussed the potential of the LIFT LLM in data-driven knowledge discovery.
Problem

Research questions and friction points this paper is trying to address.

Predicts truck driving risk using interpretable fine-tuned LLMs
Integrates literature knowledge to enhance prediction and explanation robustness
Identifies risky scenarios through variable combinations and statistical verification
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fine-tuned LLMs predict and explain truck driving risks
Literature pipeline builds knowledge base from domain papers
Evaluator assesses prediction performance and interpretability
X
Xiao Hu
Department of Civil Engineering, Tsinghua University, Beijing, China
Y
Yuansheng Lian
Department of Civil Engineering, Tsinghua University, Beijing, China
K
Ke Zhang
College of Traffic Management, People’s Public Security University of China, Beijing, China
Yunxuan Li
Yunxuan Li
Google, California Institute of Technology
PhysicsArtificial IntelligenceNatural Language Processing
Y
Yuelong Su
State Key Laboratory of Intelligent Green Vehicle and Mobility, Tsinghua University, Beijing, China
M
Meng Li
Department of Civil Engineering, Tsinghua University, Beijing, China