Watermarkable Multi-Draft Speculative Sampling via Poisson Processes

📅 2026-09-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种基于泊松过程的多稿投机采样算法,解决了大语言模型在推理效率和输出溯源间的权衡问题,同时保持了水印强度和采样效率。
📝 Abstract
Large language models (LLMs) have achieved state-of-the-art performance across a wide range of tasks, motivating two important aspects of deployment: inference efficiency and output provenance, which can be tackled by speculative sampling and watermarking, respectively. However, recent works have shown that combining these two goals is highly nontrivial and can be potentially impossible. In this work, we develop a novel multi-draft speculative sampling algorithm based on Poisson processes that improves the frontier of this fundamental trade-off. The proposed algorithm has strong sampling efficiency on its own and, more interestingly, is naturally watermarkable: we can embed an unbiased watermark without degrading speculative acceptance. Moreover, our algorithm is based on an exact list-coupling-without-communication scheme, which yields a drafter invariance property that benefits both sampling and watermarking. It is the first multi-draft, drafter-invariant speculative sampling scheme that maintains both watermark strength and sampling efficiency, and we experimentally verify its strong performance in both aspects.
Problem

Research questions and friction points this paper is trying to address.

large language models
inference efficiency
output provenance
speculative sampling
watermarking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Poisson Processes
Multi-Draft Speculative Sampling
Watermarkable
Drafter Invariance
Sampling Efficiency
🔎 Similar Papers
2023-10-27IACR Cryptology ePrint ArchiveCitations: 38
Y
Yanxiao Liu
Imperial College London
S
Sicheng Wan
University of Washington
Z
Zhan Gao
Imperial College London
D
Deniz Gündüz
Imperial College London