Checking Leakage Witnesses versus Certifying Bounded Non-Leakage

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge in language model auditing where negative (non-leakage) results are difficult to rigorously certify. By leveraging computational complexity theory and formal verification, this work systematically characterizes the complexity classes of non-leakage certification under prompt-domain constraints, complemented by Transformer architecture analysis and large-scale planted-secret experiments for quantitative evaluation. It provides the first proof that deterministic certification is coNP-complete and randomized certification is coNP^PP-complete, while demonstrating that specific attention architectures enable polynomial-time certification. Furthermore, experiments reveal a 41.06% miss rate under 256 uniform samples, delineating the coverage blind spots of finite auditing and establishing interpretive boundaries for negative audit outcomes.
📝 Abstract
When a language-model audit finds no leak, what is needed to certify non-leakage? We study guarantees over a declared prompt domain under an executable leakage criterion and decoding rule. For general bounded polynomial-time evaluators, a supplied leaking execution is polynomial-time checkable, while leak existence is \NP-complete and deterministic certification is \coNP-complete. Exact stochastic certification is $\coNP^{\PP}$-complete at every fixed rational cutoff in $(0,1)$. Restricting the computation can change these bounds. For example, certification is in \coNP\ when all randomness is a terminal draw from an efficiently computed finite probability table. Attention models admit polynomial-time certification when local dependency windows of logarithmic length precede one global head, given deterministic decoding, fixed vocabulary, exact rational weighted means, a direct binary affine readout and finite-automaton prompt domains. A construction with two global layers instead makes certification \coNP-complete over template domains, with one head per layer, polynomial width, logarithmic precision and an inverse-polynomial logit margin. Planted-secret experiments measure what finite audits miss relative to complete references. Among 30 secret--model-state pairs that leak under greedy single-prompt execution on their secret's 4,096-prompt domain, uniformly selecting 256 recorded evaluations per pair misses every leak for an expected $41.06\%$ of these pairs. Batched and single-prompt checks disagree on one complete-domain decision among all 48 fine-tuned pairs, while a same-order repeat reproduces every single-prompt output. These results distinguish computational conditions for certification from the coverage and execution conditions needed to interpret a negative audit.
Problem

Research questions and friction points this paper is trying to address.

information leakage
bounded non-leakage certification
computational complexity
language model auditing
leakage witnesses
Innovation

Methods, ideas, or system contributions that make the work stand out.

Leakage Certification
Computational Complexity
Attention Models
Language Model Auditing
Planted-Secret Experiments
🔎 Similar Papers
No similar papers found.
Chao Feng
Chao Feng
University of Zurich
networkmachine learningcybersecurity
B
Burkhard Stiller
Communication Systems Group (CSG), University of Zurich, Switzerland