natural identifier auditing

Designs and implements methods and tools to detect, measure, and localize leakage of natural identifiers from models or datasets by inserting or leveraging natural identifier canaries (random string or structured identifiers) and analyzing model outputs and behaviors. Builds audit procedures and analyses that quantify identifier exposure and support post-hoc differential-privacy style audits of trained models without requiring model retraining.

naturalidentifierauditing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.44
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of conducting efficient and scalable post-hoc privacy audits on deployed large language models, which existing methods struggle to achieve without either injecting synthetic data during training or requiring private holdout datasets drawn from the same distribution as the training data. To overcome these limitations, the paper introduces Natural Identifiers (NIDs)—structured random strings inherently present in training data, such as hash values or shortened URLs—as intrinsic signals for auditing. Leveraging NIDs, the authors develop the first general-purpose, post-training privacy auditing framework that requires no modification to the training process and no access to private reserved data. The approach supports both differential privacy verification and dataset inference, and its practicality is further enhanced through distribution-consistent synthetic data, enabling effective and scalable privacy evaluation across diverse scenarios.

dataset inferencedifferential privacylarge language models

This work addresses the risk of privacy leakage in large language models due to excessive memorization of training samples during parameter-efficient fine-tuning. To mitigate this, the authors propose a privacy auditing framework that leverages high-temperature sampling to generate interpretable, non-private, and reproducibly insertable synthetic “canary” samples, complemented by an auxiliary model auditing mechanism tailored for synthetic data. By employing membership inference and reconstruction attacks, the method precisely quantifies model memorization behavior and systematically reveals the interplay between model capacity and canary entropy in driving memorization. This approach substantially enhances the rigor and practicality of privacy auditing and establishes a novel paradigm for evaluating privacy leakage specifically in the context of synthetic data.

data leakageempirical privacy auditingmembership inference

A General Framework for Data-Use Auditing of ML Models

Jul 21, 2024
ZH
Zonghao Huang
🏛️ Duke University

To address copyright infringement and transparency concerns arising from unauthorized use of third-party data in machine learning model training, this paper proposes the first general-purpose, task-agnostic data usage auditing framework for black-box models. Methodologically, it innovatively integrates arbitrary black-box membership inference techniques with a custom sequential probability ratio test (SPRT), enabling zero assumptions about downstream tasks, strict control over false positive rates (tunable within 0.5%–5%), and cross-model generalization. The framework features a model-agnostic interface, supporting heterogeneous architectures including image classifiers and multimodal large language models. Extensive experiments on ImageNet classifiers and multimodal foundation models demonstrate an average detection accuracy exceeding 92%, with false positive rates consistently meeting user-specified thresholds. This work significantly enhances the quantifiability and reliability of training data provenance auditing.

Copyright IssuesData Usage TransparencyMachine Learning Model

To address performance degradation, reliance on white-box model information, and high false-positive rates in dataset ownership verification, this paper proposes a black-box, lossless, and zero-false-positive verification framework. Methodologically, it introduces clean-label targeted poisoning to embed a secret key—comprising out-of-distribution samples and random labels—into the training data. Post-training, the model exhibits statistically detectable, significant responses to key samples, without requiring access to internal parameters. Our key contribution is the first non-backdoor-based verification mechanism, integrating statistical hypothesis testing with ViT/ResNet ensembles. On ImageNet-1K, it achieves >99.9% detection confidence and zero accuracy loss. Moreover, it remains robust against common defenses—including pruning, fine-tuning, and input preprocessing—outperforming existing backdoor watermarking approaches significantly.

Ensures detection without harming model performanceProvides statistical certificates with black-box model accessVerifies dataset ownership via targeted data poisoning

SoK: Dataset Copyright Auditing in Machine Learning Systems

Oct 22, 2024
LD
L. Du
🏛️ Xi’an Jiaotong University | Zhejiang University | Vrije Universiteit Amsterdam | Hangzhou Dianzi University

Frequent data copyright infringement during large-scale ML model training, coupled with fragmented assumptions, narrow evaluation scopes, and poor cross-method comparability among existing copyright auditing tools, hinders practical deployment. Method: This paper systematically categorizes intrusive (watermark injection) and non-intrusive (fingerprinting-based) auditing paradigms, and—firstly—establishes a unified analytical framework spanning the entire ML pipeline: data collection, preprocessing, training, and inference. Leveraging full-stack ML modeling and controlled cross-method experiments, it characterizes structural trade-offs across assumptions, stage coverage, and real-world robustness. Contribution/Results: It introduces a deployment-oriented evaluation perspective, synthesizes common limitations, and identifies open challenges. The work delivers a taxonomy reference table and a practical implementation guide, providing both theoretical foundations and actionable technical pathways for developing compliant, deployable, and robust data copyright auditing tools.

Auditing copyright in ML training datasets to prevent unauthorized data useComparing strengths and weaknesses of existing dataset copyright auditing solutionsEvaluating robustness of auditing tools in real-world ML applications

Latest Papers

What's happening recently
View more

This work addresses the limitation of existing single-run privacy auditing methods, whose estimates of privacy leakage are often weakened due to interference among multiple canary samples. To mitigate this issue, the authors propose a novel two-level optimization strategy that integrates influence function–based greedy initialization with an embedding-space diversity objective. This approach simultaneously enhances the detectability of individual canaries and reduces their mutual interference. As a result, the method substantially strengthens the privacy leakage signal in a single training run, achieving more accurate and efficient privacy auditing than current techniques while requiring lower computational overhead.

canary craftingdifferential privacymembership inference attacks

This work addresses the risk that generative AI models may inadvertently leak private user data from their training sets through synthetic outputs, a concern exacerbated by the difficulty of distinguishing genuine privacy leakage from coincidental matches. To tackle this challenge, the paper introduces the first model-agnostic causal auditing framework that requires only synthetic outputs and a held-out reference dataset. By integrating statistical hypothesis testing with causal inference, the method rigorously differentiates true data leakage from spurious (“phantom”) matches. Applicable to any generative mechanism, the approach operates without shadow models or honeypot data, incurs computational costs orders of magnitude lower than existing techniques, and yields interpretable, tight lower bounds on privacy leakage.

data disclosuregenerative AImembership inference

This study investigates whether large language models retain poisoned identifiers in JavaScript deobfuscation even when correctly interpreting semantic meaning. Conducting 192 controlled experiments with Claude Opus 4.6 on two code prototypes, the authors employ task-specific prompt engineering, matched-pair design, and multi-step reasoning to trace identifier propagation and examine its relationship with semantic consistency. Results show that under baseline conditions, poisoned identifiers persist in 100% of outputs despite accurate semantic annotations; however, reformulating task instructions significantly reduces propagation rates to 0–20%, demonstrating the efficacy of strategic prompting. The findings highlight string-table domain consistency as a critical factor, revealing that accurate semantic understanding does not inherently lead to identifier sanitization, thereby offering new insights for security-aware code generation.

code semanticsidentifier persistencelarge language models

Existing black-box privacy scores lack interpretability because they cannot pinpoint which component within a retrieval-augmented generation (RAG) system a privacy defense actually affects. This work proposes an active path auditing method that inserts source-level hooks into the retrieval, retrieved content, and generation stages to map privacy metrics to specific leakage channels. By integrating exact-match canary testing with the NEL_strict metric to evaluate named entity leakage, the study reveals for the first time that certain differential privacy (DP)-based defenses only adjust retrieval scores without influencing generation, thereby failing to suppress named entity leakage. In contrast, end-to-end LPRAG completely blocks all 150 canary leaks in the email channel, significantly outperforming existing approaches.

black-box evaluationdefense mechanismleakage channel

This study addresses the challenge of detecting information leakage solely from a model’s predictive outputs, without access to training code, external data, or domain knowledge. Framed within decision theory, the approach models leakage diagnosis as a functional of predictive risk and outcome distribution, linking proper scoring rules with decision curve analysis via threshold-weighted associations to enable detection without prior assumptions. The work introduces a novel tripartite classification of information leakage—miscalibration, generalized calibration, and determinism—and theoretically establishes that generalized calibration leakage is fundamentally unidentifiable, whereas near-deterministic subgroups can be efficiently detected. Empirical validation on UK Biobank demonstrates detection of temporal-window comorbidity leakage down to Δc*≈0.007 in under one second, while also revealing inherent structural limitations of purely output-driven leakage detection.

blind detectiondata leakagemodel predictions