🤖 AI Summary
This study addresses the distributional discrepancy between draft and target models in speculative decoding for large language models, which constitutes both an efficiency bottleneck and a security vulnerability. We present the first investigation into the realistic attack surface of this interaction mechanism by introducing the Speculative Rejection Attack. Specifically, we design adversarial suffix generation techniques, termed Speedbump-P/D, based on probability estimation and distribution overlap optimization. These techniques elevate draft rejection rates, thereby forcing the target model to perform excessive forward passes. Our evaluation demonstrates that the proposed attack exhibits strong transferability across sampling strategies and model architectures. It degrades inference speed below that of standard autoregressive decoding while significantly increasing victim computational overhead at the cost of only marginal output quality degradation.
📝 Abstract
Speculative decoding is a popular technique for increasing the speed and reducing the costs of large language model (LLM) inference by verifying multiple draft tokens in a single target-model forward pass. The resulting benefit depends on the ability of the drafter to approximate the target model's distribution. In this work, we study Speculative Rejection Attacks (SRAs), a novel class of attacks that cause draft and target models to disagree more often, resulting in fewer draft tokens being accepted per draft cycle. This leads to more target model forward passes needed per generated token, slowing down inference and increasing costs for the victim. We introduce two attacks which append an adversarial suffix to attacker-controlled content to degrade speculative decoding on a victim's prompts. Both attacks optimise the expected length of the accepted speculative prefix, estimating per-depth acceptance from the target's probability of the drafted proposals (Speedbump-P) or from the overlap between the draft and target distributions (Speedbump-D). In some cases, attacks degrade speculative decoding to the point of being slower than autoregressive decoding. The degradation reduces the output quality - regularisation restores output quality but gives up most of the degradation, trading effectiveness for stealthiness. Additionally, the suffixes remain effective under sampling, and transfer across drafters (Speedbump-P) or across target models sharing a drafter (Speedbump-D). These findings identify the draft-target interaction of speculative decoding as a realistic attack surface through which adversarial inputs can inflate inference costs.