Controlled Decoding Attacks on Black-Box LLMs

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenges of missing probability distributions, high attack difficulty, and prohibitive sampling costs in jailbreaking black-box large language models that output text only. To this end, we propose an efficient black-box jailbreak attack framework that reconstructs the target model’s output distribution via sample distribution reconstruction. Furthermore, it introduces a novel risk-gated residual control mechanism that rebuilds distributions exclusively at critical decoding positions, substantially reducing query overhead, while incorporating speculative multi-token execution to accelerate generation. Extensive experiments conducted across four API endpoints and three benchmarks demonstrate that our approach significantly outperforms existing baselines in average scores, achieving low-cost, high-success-rate black-box jailbreak attacks.
πŸ“ Abstract
Manipulating next-token probabilities during generation can bypass the safety alignment of large language models. Existing approaches, however, rely on access to model weights or numerical token probabilities and therefore do not apply to interfaces that return only sampled text. Reconstructing probabilities from sampled outputs offers a possible alternative, but finite sampling produces sparse and noisy estimates, while repeating this process at every generation step incurs substantial query costs. Our empirical observations suggest that large distributional changes along successful jailbreak trajectories are concentrated at a small subset of positions, motivating selective control. We introduce \method{}, a framework for jailbreaking through text-only continuation interfaces that permit repeated sampling and assistant-prefix continuation. Sample-Based Distribution Reconstruction combines sampled outputs with a prior over unobserved actions to obtain a usable control signal. Risk-Gated Residual Control uses the evolving response prefix to decide when to reconstruct and modify the distribution, concentrating sampling costs at selected positions. Speculative Multi-Token Execution further amortizes target calls by verifying and accepting draft prefixes that require no intervention. Across four target endpoints and three benchmarks, \method{} achieves the highest mean score most comparisons against baselines.
Problem

Research questions and friction points this paper is trying to address.

black-box LLMs
jailbreak attack
safety alignment
controlled decoding
sample-based probability reconstruction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Black-Box Jailbreak
Distribution Reconstruction
Risk-Gated Control
Speculative Decoding
Controlled Decoding
πŸ”Ž Similar Papers