Adaptive Co-Serving LLM Watermarking on Modern Inference Engines

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing watermarking techniques for large language models are decoupled from inference engines, resulting in high latency, substantial overhead, and deployment challenges. This work proposes SWIFT, a framework that pioneers the co-design of watermark generation and inference infrastructure. By deeply integrating watermark embedding into the vLLM inference backend through instruction-guided candidate generation, asynchronous co-serving, and adaptive scheduling, SWIFT effectively eliminates redundant computation and enhances cache reuse. Experimental results demonstrate that the proposed framework achieves 99.65% detection accuracy and 97.7% robustness against adversarial attacks while preserving a text utility score of 4.87. Furthermore, SWIFT reduces end-to-end latency by 5.9× compared to baseline methods, offering a highly efficient and practical solution for deployable LLM watermarking.
📝 Abstract
Large language model (LLM) watermarking is important for ownership verification and intellectual property protection. However, existing approaches focus on algorithmic design while treating LLM inference engines as separate components. This separation often introduces auxiliary models or external tools, increasing latency and memory overhead while limiting the use of modern inference optimizations. As a result, a deployment gap remains: practical watermarking must preserve utility, detectability, and robustness while minimizing serving overhead. To bridge this gap, we propose SWIFT, a framework that co-designs LLM watermarking with modern inference infrastructure. SWIFT (i) integrates text generation and watermark construction within a shared LLM backend to reduce re-computation, improve cache reuse, and simplify system complexity; (ii) uses instruction-guided candidate generation with semantic protection to preserve factual consistency and meaning; (iii) executes watermarking concurrently with generation through asynchronous co-serving and adaptive scheduling to reduce latency and improve resource utilization; and (iv) leverages vLLM optimizations to enhance serving efficiency. Extensive experiments across domains and tasks show that SWIFT achieves strong utility, detectability, robustness, and downstream performance while substantially improving serving efficiency. Specifically, it achieves the highest utility score (4.87), 99.65% detection accuracy compared with 76.65% for the low-latency baseline SynthID, and up to 97.7% detection under watermark removal attacks, with 5.9x lower latency than the highest-utility baseline, SafeSeal. These results demonstrate the benefits of co-designing watermarking with inference infrastructure for practical LLM serving. Code and artifacts are available at: https://anonymous.4open.science/r/SWIFT-0597.
Problem

Research questions and friction points this paper is trying to address.

LLM watermarking
inference engine
serving overhead
deployment gap
latency
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM Watermarking
Co-Serving
Inference Engine Optimization
Asynchronous Scheduling
vLLM