Latent Uncertainty Representations for Video-based Driver Action and Intention Recognition

📅 2025-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing last-layer probabilistic deep learning (LL-PDL) methods for video-driven driver behavior and intention recognition under resource-constrained settings suffer from poor out-of-distribution (OOD) detection stability, inadequate calibration, and high computational overhead. Method: We propose a latent-layer uncertainty representation framework that inserts lightweight transformation layers into a pre-trained DNN to generate multi-perspective latent-space representations, coupled with repulsive training for efficient uncertainty estimation—eliminating the need for costly MCMC sampling. Contribution/Results: Our method significantly improves OOD detection performance and model calibration while preserving classification accuracy. Evaluated on four benchmark datasets, LUR/RLUR matches or exceeds state-of-the-art PDL methods in accuracy, calibration, and OOD detection, with higher training efficiency. Additionally, we contribute 28,000 frame-level action labels and 1,194 video-level intention labels to the NuScenes dataset.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationReasoning under Uncertainty: Uncertainty RepresentationsComputer Vision: Representation Learning for Vision

Application Category

Responsible Web: Machine-in-the-loop, human agency and autonomyEconomics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labelingSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
Deep neural networks (DNNs) are increasingly applied to safety-critical tasks in resource-constrained environments, such as video-based driver action and intention recognition. While last layer probabilistic deep learning (LL-PDL) methods can detect out-of-distribution (OOD) instances, their performance varies. As an alternative to last layer approaches, we propose extending pre-trained DNNs with transformation layers to produce multiple latent representations to estimate the uncertainty. We evaluate our latent uncertainty representation (LUR) and repulsively trained LUR (RLUR) approaches against eight PDL methods across four video-based driver action and intention recognition datasets, comparing classification performance, calibration, and uncertainty-based OOD detection. We also contribute 28,000 frame-level action labels and 1,194 video-level intention labels for the NuScenes dataset. Our results show that LUR and RLUR achieve comparable in-distribution classification performance to other LL-PDL approaches. For uncertainty-based OOD detection, LUR matches top-performing PDL methods while being more efficient to train and easier to tune than approaches that require Markov-Chain Monte Carlo sampling or repulsive training procedures.
Problem

Research questions and friction points this paper is trying to address.

Improving uncertainty estimation in video-based driver action recognition
Enhancing out-of-distribution detection for safety-critical autonomous systems
Developing efficient uncertainty methods without complex training procedures
Innovation

Methods, ideas, or system contributions that make the work stand out.

Extending pre-trained DNNs with transformation layers
Producing multiple latent representations for uncertainty estimation
Achieving efficient OOD detection without complex training procedures
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Koen Vellenga
Koen Vellenga
Unknown affiliation
deep learninguncertaintyintention recognition
H
H. Joe Steinhauer
University of Skövde, Sweden
J
Jonas Andersson
Volvo Car Corporation, Sweden
A
Anders Sjögren
Volvo Car Corporation, Sweden