🤖 AI Summary
This study addresses a security vulnerability in length-prediction-based scheduling strategies for large language model (LLM) serving, wherein adversaries can manipulate input signals to obtain high scheduling priority. We are the first to reveal significant discrepancies between predicted lengths and actual outputs, and propose JIL, an adversarial attack method that optimizes adversarial suffixes to mislead lightweight length probe models, thereby deliberately reducing predicted values to circumvent scheduling fairness. Experimental results demonstrate that this attack reduces predicted lengths by up to 83.4% and accelerates request processing by an average factor of 1.53×. Furthermore, we validate that coarse-grained grouping defense mechanisms can effectively mitigate this threat.
📝 Abstract
Efficient request scheduling is increasingly important for reducing completion time in large language model (LLM) serving. Size-based policies such as Shortest Job First prioritize shorter requests, but output lengths are unknown before generation, so practical schedulers rely on predicted lengths. We introduce JIL, an attack on prediction-based LLM schedulers that manipulates the scheduling signal to obtain higher priority and reduce completion time. Using TRAIL as a case study, JIL optimizes an adversarial suffix that causes a lightweight output-length probe to underestimate a request's length. We evaluate JIL on two datasets and four LLMs across varied request profiles and deployment configurations. JIL reduces predicted output lengths by up to 83.4 percent, and adversarial requests complete up to 1.53 times faster on average in end-to-end serving experiments. The reduction in predicted length is substantially larger than the change in actual output length, revealing a mismatch between the scheduler's estimate and the request's realized size. Response utility varies across models and tasks, exposing a trade-off between scheduling advantage and response quality. We also evaluate scheduler-side defenses and find that grouping length predictions into coarse intervals reduces JIL's scheduling advantage and mitigates delays to benign requests.