Practical Secrets Extraction against Black-box LLMs

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the risk of confidential credential memorization in black-box large language models due to training data leakage, as well as the limitations of existing security audits that rely solely on output access. We propose a pioneering output-only secret extraction framework for black-box models. This approach constructs surrogate models via cross-validated knowledge distillation and incorporates provider-specific structural priors. It further integrates techniques including semantics-preserving prompt variations, truncated top-p sampling, local token entropy analysis, and n-gram frequency profiling to achieve efficient extraction. Experimental results demonstrate that the proposed framework significantly outperforms mainstream methods across benchmarks, successfully recovering masked credentials from real-world deployed systems such as OpenAI and Claude. These findings expose critical security vulnerabilities inherent in current black-box model deployments.
📝 Abstract
Large language models (LLMs) increasingly power autonomous coding agents such as Codex and Claude Code, yet their training corpora may contain confidential credentials exposed in public repositories or collected from private development artifacts, creating risks of memorization and subsequent leakage. Existing extraction audits, however, largely assume access to model weights or token probabilities. In this work, we present a black-box secret extraction framework for commercial, API-based LLMs under output-only access. It comprises (i) \emph{Cross-Validated Secret Knowledge Distillation}, which uses semantics-preserving prompt variants, response cross-validation, and provider-specific format filtering to distill secret-relevant behavior into a local white-box proxy; and (ii) \emph{Proxy-Guided Secret Extraction and Candidate Filtering}, which combines truncated top-$p$ sampling with local token entropy, $N$-gram frequency profiling, and provider-specific structural priors. On controlled API-key benchmarks, our framework improves recovery effectiveness and real-key rates over representative baselines while reducing extraction latency. A responsible real-world evaluation further recovers masked provider-specific credentials from three independently deployed black-box LLM systems spanning OpenAI and Claude Code, showing that memorized secrets can be exposed under output-only access.
Problem

Research questions and friction points this paper is trying to address.

Black-box LLMs
Secret Extraction
Credential Leakage
Memorization
Output-only Access
Innovation

Methods, ideas, or system contributions that make the work stand out.

Black-box Extraction
Knowledge Distillation
Secret Recovery
Proxy-Guided Sampling
LLM Security
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Shiqian Zhao
Shiqian Zhao
Nanyang Technological University of Singapore
RobustAIAI SecurityAutomatic Driving
S
Siwei Jiang
Beijing University of Posts and Telecommunications
X
Xinfeng Li
Hong Kong Polytechnic University
Runyi Hu
Runyi Hu
Nanyang Technological University
Large Language ModelAI AlignmentWatermarking
Y
Yandan Zheng
Nanyang Technological University
C
Congyu Guo
Beijing University of Posts and Telecommunications
Tianwei Zhang
Tianwei Zhang
Nanyang Technological University
Computer System Security
A
Anh Tuan Luu
Nanyang Technological University