implement data masking

Designs, implements, and evaluates transformations and tooling that conceal, replace, or obfuscate sensitive values in datasets or data streams—using techniques such as redaction, tokenization, perturbation, encryption, and observation/masked modeling—to enable safe downstream use. Builds these masking processes into data pipelines, storage, and APIs and measures utility-versus-privacy trade-offs and residual re-identification risk.

implementdatamasking

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$199K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

High-sensitivity sectors such as Banking, Financial Services, and Insurance (BFSI) face a fundamental trade-off between privacy protection and analytical utility when applying large-scale data analytics. Method: This paper proposes a multi-layered, privacy-enhancing framework that synergistically integrates conditional generative adversarial networks (cGANs) for high-fidelity synthetic data generation, context-aware PII transformation, configurable statistical perturbation, and differential privacy mechanisms—collectively optimizing the privacy–utility Pareto frontier. Contribution/Results: The resulting end-to-end, configurable pipeline replaces conventional anonymization techniques and achieves, in real-world BFSI deployments: <3% model training accuracy degradation, 92% reduction in privacy leakage risk, and 40% acceleration in analytical cycle time—fully complying with GDPR and CCPA requirements. To our knowledge, this is the first synthetic data solution for high-sensitivity domains that simultaneously delivers strong privacy guarantees, high statistical fidelity, and production-grade deployability.

Balancing privacy and utility in synthetic data generationComparing modern perturbation techniques with traditional anonymizationEnhancing security and efficiency in BFSI data management

This work systematically uncovers a critical privacy risk in existing dataset distillation methods: when compressing real data into synthetic data, these approaches may implicitly encode the training trajectory of models, thereby leaking sensitive information about the original dataset. To expose this vulnerability, the authors propose an Information Revelation Attack (IRA) that integrates model inversion and membership inference techniques to effectively infer the distillation algorithm, model architecture, membership status, and even reconstruct sensitive samples from the synthetic data alone. Experimental results demonstrate that IRA can accurately identify both the distillation method and model structure, and successfully recover original sensitive data with high fidelity. These findings fundamentally challenge the prevailing assumption that dataset distillation inherently preserves privacy, revealing instead a severe and previously underappreciated privacy leakage risk.

dataset distillationinformation revelationmembership inference

This paper presents the first systematic evaluation of commercial large language models (LLMs) on assembly code deobfuscation. Addressing four prevalent obfuscation techniques—control-flow flattening, bogus control flow, instruction substitution, and their combinations—the authors propose a four-dimensional theoretical framework (reasoning depth, pattern recognition, noise filtering, and contextual integration) and develop a three-tier obfuscation resistance model. Through prompt engineering and semantic modeling, they conduct empirical analysis across diverse obfuscation scenarios. Results show that LLMs autonomously handle low-resistance obfuscations (e.g., bogus control flow) but completely fail on composite obfuscations. The study identifies five canonical error patterns, exposing fundamental limitations in deep semantic comprehension and structural recovery. Collectively, this work provides both theoretical characterization and empirical evidence delineating the applicability boundaries of LLMs in reverse engineering tasks.

Evaluating LLMs' ability to deobfuscate assembly codeIdentifying performance variations across obfuscation scenariosProposing a framework to explain LLM deobfuscation limitations

To address performance degradation, reliance on white-box model information, and high false-positive rates in dataset ownership verification, this paper proposes a black-box, lossless, and zero-false-positive verification framework. Methodologically, it introduces clean-label targeted poisoning to embed a secret key—comprising out-of-distribution samples and random labels—into the training data. Post-training, the model exhibits statistically detectable, significant responses to key samples, without requiring access to internal parameters. Our key contribution is the first non-backdoor-based verification mechanism, integrating statistical hypothesis testing with ViT/ResNet ensembles. On ImageNet-1K, it achieves >99.9% detection confidence and zero accuracy loss. Moreover, it remains robust against common defenses—including pruning, fine-tuning, and input preprocessing—outperforming existing backdoor watermarking approaches significantly.

Ensures detection without harming model performanceProvides statistical certificates with black-box model accessVerifies dataset ownership via targeted data poisoning

A False Sense of Privacy: Evaluating Textual Data Sanitization Beyond Surface-level Privacy Leakage

Apr 28, 2025
RX
Rui Xin
🏛️ University of Washington | Allen Institute for Artificial Intelligence

This paper identifies a pervasive “privacy illusion” in text anonymization: existing methods—such as PII removal or synthetic data generation—only mitigate explicit identifier leakage, yet remain vulnerable to semantic re-identification attacks. To address this, we propose the first risk assessment framework for semantic-level privacy leakage, integrating empirical re-identification attacks, benchmarking against mainstream commercial PII detection tools (e.g., Azure), differential privacy baselines, and evaluation on the MedQA medical QA dataset. Results reveal that Azure fails to protect 74% of sensitive information in MedQA; while differential privacy reduces re-identification risk, it severely degrades textual utility. Our work challenges prevailing assumptions about anonymization efficacy and establishes a verifiable, empirically grounded evaluation paradigm for semantic privacy protection.

Assessing re-identification attacks using nuanced textual markers for privacy breachesBalancing privacy protection and data utility in text sanitization techniquesEvaluating privacy risks beyond explicit identifiers in sanitized text data

Latest Papers

What's happening recently
View more

Existing image protection methods exhibit insufficient robustness against heterogeneous diffusion model attacks and lack systematic evaluation under model mismatch scenarios. This work proposes a unified post-publication purification framework that restores image editability without requiring the original image or internal knowledge of the defense mechanism, even when the attacker’s and defender’s model architectures differ. We are the first to systematically uncover failure modes of the “purify once, edit freely” paradigm and introduce two practical purifiers: a VAE-Transformer-based latent space projection correction module and EditorClean, a Diffusion Transformer-driven instruction-guided reconstruction model. Evaluated across 2,100 editing tasks, EditorClean improves PSNR by 3–6 dB and reduces FID by 50–70% over existing baselines, demonstrating that most protection mechanisms fail after purification.

adversarial perturbationsdiffusion modelsimage protection

This study addresses a critical gap in synthetic data generation (SDG) research, which has predominantly focused on privacy attacks initiated by data recipients while overlooking internal adversaries—such as data owners or generators—who may degrade data quality by perturbing real data. The work formally introduces this internal threat model and proposes targeted perturbation strategies based on label flipping and feature importance manipulation. Through systematic evaluation across multiple mainstream SDG frameworks, the experiments demonstrate that even minimal perturbations can substantially impair downstream task performance and amplify statistical distributional biases. These findings reveal a pronounced vulnerability in current SDG pipelines regarding data integrity and underscore the urgent need for robustness and integrity verification mechanisms in synthetic data workflows.

Adversarial ManipulationData IntegrityPrivacy-Preserving Data Sharing

Bloom Filter Encoding for Machine Learning

Dec 22, 2025
JC
John Cartmell
🏛️ Florida Atlantic University

This paper addresses three key challenges in machine learning data preprocessing: high memory overhead, significant privacy leakage risk, and difficulty preserving structural information. To tackle these, we propose a novel Bloom filter–based encoding method that maps raw samples into compact, irreversible binary vectors. This work is the first to systematically validate Bloom filters as a general-purpose, privacy-enhancing feature transformation technique. Empirical evaluation across six heterogeneous datasets—using XGBoost, DNNs, CNNs, and logistic regression—demonstrates classification accuracy comparable to that achieved on raw data (average degradation <1.2%), while reducing memory consumption by up to 87%. Crucially, the original features are provably unrecoverable, thereby eliminating reconstruction-based privacy threats. Our core contribution lies in establishing both the theoretical applicability and empirical superiority of Bloom filters for lightweight, privacy-preserving preprocessing.

Encodes data into privacy-preserving bit arraysProvides efficient preprocessing for diverse machine learning tasksReduces memory usage while maintaining classification accuracy

This work addresses the privacy risks posed by sensitive content in large-scale image datasets by proposing a two-stage automated anonymization framework. First, a vision-language model identifies privacy-sensitive regions and generates paired private/public textual descriptions along with structured editing instructions. Subsequently, an instruction-driven diffusion editor precisely rewrites sensitive visual content to preserve semantic integrity while ensuring privacy. This study is the first to integrate multimodal guidance with structured instructions for controllable anonymization and introduces a unified evaluation framework encompassing privacy preservation, fidelity, and downstream task utility. Experimental results demonstrate that the method significantly reduces facial similarity, textual identifiability, and demographic predictability, while maintaining downstream task performance comparable to that of the original data.

dataset safetydownstream utilityimage anonymization

This work identifies and formally names a previously unrecognized privacy vulnerability in mainstream machine learning frameworks, termed “Quantamination,” arising from the implementation of dynamic quantization. While dynamic quantization enhances inference efficiency, we demonstrate that it inadvertently introduces a novel side-channel attack surface, enabling cross-batch leakage of user inputs. Through a systematic combination of side-channel analysis, reverse engineering of quantization mechanisms, and auditing of framework configurations, we empirically evaluate multiple widely used ML inference engines. Our findings confirm that at least four frameworks, under default or common deployment settings, are susceptible to this vulnerability, allowing an adversary to partially or fully reconstruct sensitive data from other users within the same inference batch.

batch privacydata leakagedynamic quantization

Hot Scholars

YC

Yujun Cai

NTU → Meta → Lecturer(Assistant Professor) @UQ
Multi-Modal PerceptionVision-Language Models
YC

Yu-Chiang Frank Wang

National Taiwan University & NVIDIA
Computer VisionDeep LearningMachine LearningArtificial Intelligence
GL

Gaowen Liu

Cisco Research
machine learningcomputer visionmultimedia.
HD

Hao Deng

Engineer
recommendation system