Deep Reinforcement Learning for Misbehavior Detection Under Partially Observable V2X Data

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of misbehavior detection in V2X environments, where partial data observability enables adversaries to evade detection by exploiting missing information. We propose an adaptive detection framework based on deep reinforcement learning (DRL) that learns dynamic policies over incomplete V2X data streams. To evaluate robustness, we construct a threat model integrating natural occlusion with adversarial feature suppression. This work pioneers the application of DRL to partially observable V2X scenarios, exposing the vulnerability of static baselines under adversarial data absence and establishing a novel attack–defense evaluation paradigm. Experimental results demonstrate that the proposed DRL approach significantly outperforms an XGBoost baseline under natural partial observability; however, it remains vulnerable to strategic data withholding, highlighting critical directions for future resilient system design.
📝 Abstract
Misbehavior detection in vehicle-to-everything (V2X) systems is essential for ensuring the semantic correctness of exchanged messages and preventing the dissemination of falsified information. Existing data-centric misbehavior detection approaches largely rely on statistical validation or supervised machine learning models under the implicit assumption of fully observable V2X streams. In practice, however, vehicular environments are inherently partially observable due to hardware failures, intermittent connectivity, and environmental occlusions. Moreover, missingness itself can be strategically exploited by adversaries to evade detection. In this paper, we study misbehavior detection under incomplete V2X observations and propose a deep reinforcement learning (DRL)-based detection framework that learns adaptive policies with incomplete data. We further introduce an adversarial threat model in which attackers exploit or deliberately induce missingness to evade detection, including evasion via natural occlusions and adversarial feature suppression. Extensive experiments conducted on the VeReMi dataset under various missingness patterns demonstrate that DRL significantly outperforms a powerful XGBoost baseline under natural partial observability. However, results also reveal a critical vulnerability: DRL policies can be highly susceptible to evasion attacks that strategically exploit natural missingness. In contrast, DRL exhibits more gradual degradation under direct feature suppression compared to static tree-based models.
Problem

Research questions and friction points this paper is trying to address.

Misbehavior Detection
V2X
Partial Observability
Adversarial Evasion
Incomplete Data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Deep Reinforcement Learning
Misbehavior Detection
Partially Observable V2X
Adversarial Threat Model
Evasion Attacks
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Roshan Sedar
Sustainable Artificial Intelligence Research Unit, Centre Tecnològic de Telecomunicacions de Catalunya (CTTC/CERCA), 08860 Castelldefels, Spain
Charalampos Kalalas
Charalampos Kalalas
Centre Tecnològic de Telecomunicacions de Catalunya (CTTC)
probabilistic modellingcomputational statisticsresource-efficient MLanomaly detectionwireless