Surviving the Unseen: Predictive Defense for Novel Multi-Turn Multimodal Attacks

📅 2026-05-18
📈 Citations: 0
Influential: 0
📄 PDF

career value

196K/year
🤖 AI Summary
Existing static defense mechanisms struggle to counter progressive adversarial attacks that span multiple turns and modalities, as they overlook the cumulative structural contamination embedded in dialogue trajectories. This work proposes TRIAD, a novel framework that uniquely integrates structural anomaly detection, Ledoit-Wolf regularized Mahalanobis distance, and topological trajectory acceleration within a time-varying Cox proportional hazards model augmented by Bayesian hidden Markov feedback. By formulating security verification as a dynamic survival prediction problem, TRIAD enables early detection and interception of malicious intent without requiring model retraining. The approach provides a theoretical upper bound on failure time under adversarial perturbations, ensuring accelerated divergence of malicious trajectories. Consequently, it offers multimodal conversational agents an efficient, interpretable, and theoretically grounded mechanism for continuous safety alignment.
📝 Abstract
The expansion of Multimodal Large Language Models (MLLMs) and their integration into autonomous agentic workflows has introduced a non-stationary attack surface. Empirical observations indicate that adversaries employ progressive, cross-modal perturbations that evade turn-specific guardrails by distributing malicious intent across longitudinal conversational trajectories. Static defense mechanisms, constrained by the Markov property, evaluate inputs in isolation and fail to detect cumulative structural poisoning. To handle this limitation, this paper formulates safety verification as a dynamic survival prediction and trajectory dynamics problem. The Triple-tier Anomaly Defense (TRIAD) framework is proposed as a predictive model that maps multimodal and multi-turn conversational flow as a continuous trajectory. The framework integrates structural anomaly detection to monitor covariance shifts, a Ledoit-Wolf regularized Mahalanobis distance to monitor covariance shifts in high-dimensional spaces, and topological trajectory acceleration to differentiate benign creative exploration from continuous malicious drift. These kinematic and geometric features are integrated into a time-varying Cox Proportional Hazards model via a Bayesian Hidden Markov Model (HMM) feedback loop. Theoretical analysis demonstrates that the TRIAD framework provides a mathematically bounded expected time-to-failure under adversarial perturbations, ensuring that malicious acceleration diverges positively. This framework provides a computationally efficient, interpretable, and predictive safeguard for real-time agentic AI systems, establishing a rigorous foundation for continuous safety alignment without relying on empirical retraining.
Problem

Research questions and friction points this paper is trying to address.

multimodal attacks
multi-turn conversations
adversarial perturbations
structural poisoning
non-stationary attack surface
Innovation

Methods, ideas, or system contributions that make the work stand out.

predictive defense
multimodal attacks
trajectory dynamics
anomaly detection
survival analysis
🔎 Similar Papers
No similar papers found.