Deep Learning Latency Attacks and Defenses: A Cross-Domain Survey of Availability Threats

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses availability threats to deep learning systems, such as response timeouts and resource exhaustion induced by latency attacks. Conducting a cross-domain survey spanning perception, inference, and agent systems, it transcends conventional attack-form taxonomies by unifying fragmented literature through the lens of computational bottlenecks. The study proposes a defense abstraction based on work budgets alongside a system-level failure analysis framework, integrating techniques from adversarial machine learning, adaptive neural inference, and large model security for in-depth examination. Ultimately, it establishes a threat model taxonomy, provides quantitative comparisons and a comprehensive review of defense strategies, and delineates open challenges for future research.
📝 Abstract
Adversarial machine learning has focused mainly on integrity, but availability is an increasingly consequential complement. Latency attacks (also energy-latency attacks) increase inference-time work, energy, or response time, causing deadline misses, throughput collapse, or resource exhaustion in vehicle controllers, interactive services, or battery-powered sensors, sometimes while preserving the nominal prediction. This survey unifies a fragmented literature spanning perception pipelines (including physical attacks on autonomous-driving detection and tracking), input-adaptive neural inference (sponge examples, dynamic networks), and autoregressive and agentic systems (output-length, verbose-image, and reasoning denial-of-service attacks on LLMs, VLMs, mixture-of-experts models, and tool-using agents). We organize attacks by exploited computational bottleneck rather than formulation, separating what makes a computation expensive from how the attacker triggers it; the delivery channel (input, prompt or retrieved content, message, poisoning, or weight tampering) is an orthogonal attribute. Many attacks share one mechanism, intermediate-work amplification, motivating a work-budget defense abstraction; we distinguish caps on the work entering an expensive stage from caps on the results leaving it. We further analyze when a model-level cost increase becomes a system-level availability failure, which depends on critical-path share, slack, existing ceilings, accumulation, resource sharing, and fallback policy, not on the amplification factor alone. We also provide a threat-model taxonomy, consolidated quantitative comparisons, a defense review by control mechanism, and open challenges such as standardized evaluation, physical realizability, and whole-system availability. Companion website: https://github.com/guzonghua/awesome-latency-attacks.
Problem

Research questions and friction points this paper is trying to address.

Latency attacks
Deep learning availability
Adversarial machine learning
Denial-of-service
Inference-time resource exhaustion
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latency Attacks
Intermediate-Work Amplification
Work-Budget Defense
System-Level Availability
Computational Bottleneck
🔎 Similar Papers
No similar papers found.