Score
Designs, implements, and evaluates transformations and tooling that conceal, replace, or obfuscate sensitive values in datasets or data streams—using techniques such as redaction, tokenization, perturbation, encryption, and observation/masked modeling—to enable safe downstream use. Builds these masking processes into data pipelines, storage, and APIs and measures utility-versus-privacy trade-offs and residual re-identification risk.
High-sensitivity sectors such as Banking, Financial Services, and Insurance (BFSI) face a fundamental trade-off between privacy protection and analytical utility when applying large-scale data analytics. Method: This paper proposes a multi-layered, privacy-enhancing framework that synergistically integrates conditional generative adversarial networks (cGANs) for high-fidelity synthetic data generation, context-aware PII transformation, configurable statistical perturbation, and differential privacy mechanisms—collectively optimizing the privacy–utility Pareto frontier. Contribution/Results: The resulting end-to-end, configurable pipeline replaces conventional anonymization techniques and achieves, in real-world BFSI deployments: <3% model training accuracy degradation, 92% reduction in privacy leakage risk, and 40% acceleration in analytical cycle time—fully complying with GDPR and CCPA requirements. To our knowledge, this is the first synthetic data solution for high-sensitivity domains that simultaneously delivers strong privacy guarantees, high statistical fidelity, and production-grade deployability.
This work systematically uncovers a critical privacy risk in existing dataset distillation methods: when compressing real data into synthetic data, these approaches may implicitly encode the training trajectory of models, thereby leaking sensitive information about the original dataset. To expose this vulnerability, the authors propose an Information Revelation Attack (IRA) that integrates model inversion and membership inference techniques to effectively infer the distillation algorithm, model architecture, membership status, and even reconstruct sensitive samples from the synthetic data alone. Experimental results demonstrate that IRA can accurately identify both the distillation method and model structure, and successfully recover original sensitive data with high fidelity. These findings fundamentally challenge the prevailing assumption that dataset distillation inherently preserves privacy, revealing instead a severe and previously underappreciated privacy leakage risk.
This paper presents the first systematic evaluation of commercial large language models (LLMs) on assembly code deobfuscation. Addressing four prevalent obfuscation techniques—control-flow flattening, bogus control flow, instruction substitution, and their combinations—the authors propose a four-dimensional theoretical framework (reasoning depth, pattern recognition, noise filtering, and contextual integration) and develop a three-tier obfuscation resistance model. Through prompt engineering and semantic modeling, they conduct empirical analysis across diverse obfuscation scenarios. Results show that LLMs autonomously handle low-resistance obfuscations (e.g., bogus control flow) but completely fail on composite obfuscations. The study identifies five canonical error patterns, exposing fundamental limitations in deep semantic comprehension and structural recovery. Collectively, this work provides both theoretical characterization and empirical evidence delineating the applicability boundaries of LLMs in reverse engineering tasks.
To address performance degradation, reliance on white-box model information, and high false-positive rates in dataset ownership verification, this paper proposes a black-box, lossless, and zero-false-positive verification framework. Methodologically, it introduces clean-label targeted poisoning to embed a secret key—comprising out-of-distribution samples and random labels—into the training data. Post-training, the model exhibits statistically detectable, significant responses to key samples, without requiring access to internal parameters. Our key contribution is the first non-backdoor-based verification mechanism, integrating statistical hypothesis testing with ViT/ResNet ensembles. On ImageNet-1K, it achieves >99.9% detection confidence and zero accuracy loss. Moreover, it remains robust against common defenses—including pruning, fine-tuning, and input preprocessing—outperforming existing backdoor watermarking approaches significantly.
This paper identifies a pervasive “privacy illusion” in text anonymization: existing methods—such as PII removal or synthetic data generation—only mitigate explicit identifier leakage, yet remain vulnerable to semantic re-identification attacks. To address this, we propose the first risk assessment framework for semantic-level privacy leakage, integrating empirical re-identification attacks, benchmarking against mainstream commercial PII detection tools (e.g., Azure), differential privacy baselines, and evaluation on the MedQA medical QA dataset. Results reveal that Azure fails to protect 74% of sensitive information in MedQA; while differential privacy reduces re-identification risk, it severely degrades textual utility. Our work challenges prevailing assumptions about anonymization efficacy and establishes a verifiable, empirically grounded evaluation paradigm for semantic privacy protection.
Existing image protection methods exhibit insufficient robustness against heterogeneous diffusion model attacks and lack systematic evaluation under model mismatch scenarios. This work proposes a unified post-publication purification framework that restores image editability without requiring the original image or internal knowledge of the defense mechanism, even when the attacker’s and defender’s model architectures differ. We are the first to systematically uncover failure modes of the “purify once, edit freely” paradigm and introduce two practical purifiers: a VAE-Transformer-based latent space projection correction module and EditorClean, a Diffusion Transformer-driven instruction-guided reconstruction model. Evaluated across 2,100 editing tasks, EditorClean improves PSNR by 3–6 dB and reduces FID by 50–70% over existing baselines, demonstrating that most protection mechanisms fail after purification.
This study addresses a critical gap in synthetic data generation (SDG) research, which has predominantly focused on privacy attacks initiated by data recipients while overlooking internal adversaries—such as data owners or generators—who may degrade data quality by perturbing real data. The work formally introduces this internal threat model and proposes targeted perturbation strategies based on label flipping and feature importance manipulation. Through systematic evaluation across multiple mainstream SDG frameworks, the experiments demonstrate that even minimal perturbations can substantially impair downstream task performance and amplify statistical distributional biases. These findings reveal a pronounced vulnerability in current SDG pipelines regarding data integrity and underscore the urgent need for robustness and integrity verification mechanisms in synthetic data workflows.
This paper addresses three key challenges in machine learning data preprocessing: high memory overhead, significant privacy leakage risk, and difficulty preserving structural information. To tackle these, we propose a novel Bloom filter–based encoding method that maps raw samples into compact, irreversible binary vectors. This work is the first to systematically validate Bloom filters as a general-purpose, privacy-enhancing feature transformation technique. Empirical evaluation across six heterogeneous datasets—using XGBoost, DNNs, CNNs, and logistic regression—demonstrates classification accuracy comparable to that achieved on raw data (average degradation <1.2%), while reducing memory consumption by up to 87%. Crucially, the original features are provably unrecoverable, thereby eliminating reconstruction-based privacy threats. Our core contribution lies in establishing both the theoretical applicability and empirical superiority of Bloom filters for lightweight, privacy-preserving preprocessing.
This work addresses the privacy risks posed by sensitive content in large-scale image datasets by proposing a two-stage automated anonymization framework. First, a vision-language model identifies privacy-sensitive regions and generates paired private/public textual descriptions along with structured editing instructions. Subsequently, an instruction-driven diffusion editor precisely rewrites sensitive visual content to preserve semantic integrity while ensuring privacy. This study is the first to integrate multimodal guidance with structured instructions for controllable anonymization and introduces a unified evaluation framework encompassing privacy preservation, fidelity, and downstream task utility. Experimental results demonstrate that the method significantly reduces facial similarity, textual identifiability, and demographic predictability, while maintaining downstream task performance comparable to that of the original data.
This work identifies and formally names a previously unrecognized privacy vulnerability in mainstream machine learning frameworks, termed “Quantamination,” arising from the implementation of dynamic quantization. While dynamic quantization enhances inference efficiency, we demonstrate that it inadvertently introduces a novel side-channel attack surface, enabling cross-batch leakage of user inputs. Through a systematic combination of side-channel analysis, reverse engineering of quantization mechanisms, and auditing of framework configurations, we empirically evaluate multiple widely used ML inference engines. Our findings confirm that at least four frameworks, under default or common deployment settings, are susceptible to this vulnerability, allowing an adversary to partially or fully reconstruct sensitive data from other users within the same inference batch.