Cooldown Landmines: Cross-Tenant Interference Attacks on LLM Gateways

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses cross-tenant interference in LLM gateways caused by shared cooldown records, which adversaries can exploit via fault-handling mechanisms to restrict other tenants from accessing available backends. Specifically, this work identifies two cross-tenant attacks leveraging shared cooldown states and proposes source verification alongside tenant-level cooldown mechanisms to isolate fault impacts. To mitigate RPM-based denial attacks, a source verification strategy is introduced, while tenant-level cooldowns are achieved by retaining shared backend failure records, thereby bridging the isolation gap between request admission and fault handling. Experimental results demonstrate that the proposed approach effectively defends against forced fallback and denial-of-service attacks, reducing the post-recovery victim fallback rate from 54/60 to 0/60.
📝 Abstract
LLM gateways enforce separate tenant quotas while sharing model deployments and cooldown records that temporarily exclude failing backends. However, a tenant's request failure can update these shared records and restrict other tenants'access to serviceable deployments. We identify two attacks that exploit this gap in LiteLLM. The first uses requests rejected at the key's requests-per-minute (RPM) limit: caller-supplied identifiers for known registered deployments reach failure handling, allowing two rejected requests to redirect another tenant to fallback with zero upstream calls from those requests. The second uses admitted traffic to create cooldown records that persist after backend capacity recovers. To address these failures, we design an origin check that blocks deployment updates from key RPM rejections and tenant-scoped cooldown that preserves the triggering tenant's back-off while retaining shared records for backend faults. Experiments with authenticated proxies and self-hosted vLLM demonstrate that the attacks can force fallback or denial while deployments remain serviceable. Across five paired four-worker runs, tenant scoping reduces victim fallback after recovery from 54/60 to 0/60, while increasing attempts against exhausted shared quotas. These findings show that tenant isolation must cover both request admission and the failure handling that governs shared deployment availability.
Problem

Research questions and friction points this paper is trying to address.

LLM gateways
cross-tenant interference
cooldown records
denial of service
tenant isolation
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-tenant interference
LLM gateway security
cooldown mechanism
tenant isolation
rate limiting attack
🔎 Similar Papers
No similar papers found.