sanitize outputs

Designs and implements systems and postprocessing pipelines that detect, remove, or redact unsafe, private, or sensitive information from model outputs and related artifacts, including content sanitization, sensitive-data scrubbing, redaction of absolute file paths and PDK/model identifiers, and masking of license-bound state and secrets. Builds filters and transformation rules that produce safe alternatives (e.g., numeric summaries, intent-only responses, reframed content) while preserving as much legitimate usefulness as possible under privacy and safety constraints.

sanitizeoutputs

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.52
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the critical security and ethical risks arising from enterprise employees inadvertently leaking sensitive data or generating policy-violating, unethical content when using large language models. To mitigate these risks, we propose SafeGPT—the first unified dual-sided protection framework that integrates input-side sensitive information detection and sanitization with output-side content moderation and rewriting. SafeGPT further incorporates a human-in-the-loop feedback mechanism to jointly optimize safety and user experience. By combining red-teaming attacks with reinforcement learning from human feedback, the system significantly reduces the likelihood of data leakage and biased outputs while maintaining high user satisfaction.

data leakageenterprise LLMsethics

This work addresses the privacy risks posed by sensitive content in large-scale image datasets by proposing a two-stage automated anonymization framework. First, a vision-language model identifies privacy-sensitive regions and generates paired private/public textual descriptions along with structured editing instructions. Subsequently, an instruction-driven diffusion editor precisely rewrites sensitive visual content to preserve semantic integrity while ensuring privacy. This study is the first to integrate multimodal guidance with structured instructions for controllable anonymization and introduces a unified evaluation framework encompassing privacy preservation, fidelity, and downstream task utility. Experimental results demonstrate that the method significantly reduces facial similarity, textual identifiability, and demographic predictability, while maintaining downstream task performance comparable to that of the original data.

dataset safetydownstream utilityimage anonymization

Truthful Text Sanitization Guided by Inference Attacks

Dec 17, 2024
IP
Ildik'o Pil'an
🏛️ Norwegian Computing Center | Universitat Rovira i Virgili

Text anonymization must balance privacy preservation with semantic utility. This paper proposes an automated anonymization method based on abstraction-based generalization: first, instruction-tuned large language models (LLMs) generate fidelity-preserving substitution candidates and rank them by abstraction level; second, an LLM-driven inference attack simulation quantifies each candidate’s resistance to re-identification; finally, authenticity, abstraction, and privacy robustness are jointly optimized to select the optimal substitution. Key contributions include: (i) the first integration of inference attacks into the anonymization decision loop; (ii) a novel, annotation-free metric jointly evaluating utility and privacy; and (iii) end-to-end multi-objective co-optimization. On the Text Anonymization Benchmark, our method achieves significantly higher utility than baselines, incurs only marginally higher re-identification risk than full suppression, and yields substitutions with superior fidelity and abstraction.

Balancing privacy protection and content utility in text sanitizationDeveloping truth-preserving replacements resistant to inference attacksPreventing personal information leakage while preserving document semantics

A False Sense of Privacy: Evaluating Textual Data Sanitization Beyond Surface-level Privacy Leakage

Apr 28, 2025
RX
Rui Xin
🏛️ University of Washington | Allen Institute for Artificial Intelligence

This paper identifies a pervasive “privacy illusion” in text anonymization: existing methods—such as PII removal or synthetic data generation—only mitigate explicit identifier leakage, yet remain vulnerable to semantic re-identification attacks. To address this, we propose the first risk assessment framework for semantic-level privacy leakage, integrating empirical re-identification attacks, benchmarking against mainstream commercial PII detection tools (e.g., Azure), differential privacy baselines, and evaluation on the MedQA medical QA dataset. Results reveal that Azure fails to protect 74% of sensitive information in MedQA; while differential privacy reduces re-identification risk, it severely degrades textual utility. Our work challenges prevailing assumptions about anonymization efficacy and establishes a verifiable, empirically grounded evaluation paradigm for semantic privacy protection.

Assessing re-identification attacks using nuanced textual markers for privacy breachesBalancing privacy protection and data utility in text sanitization techniquesEvaluating privacy risks beyond explicit identifiers in sanitized text data

Latest Papers

What's happening recently
View more

This work addresses the risk that agent execution trajectories may inadvertently leak proprietary procedural knowledge—such as critical formulas, thresholds, and decision strategies—even when model weights remain undisclosed. To mitigate this, the authors propose RedAct, a novel framework that treats execution trajectories as security-sensitive interfaces. RedAct integrates sensitive information localization, semantics-preserving trajectory rewriting, and behavioral watermark embedding to sanitize trajectories while retaining essential evidence required for auditability. Evaluated on the newly introduced CapTraceBench benchmark, RedAct reduces normalized skill transfer rates in diverse trajectory reuse scenarios to below the no-skill baseline (originally 44.7–67.1%), achieves behavioral watermark detection rates of 93.6–100.0%, and maintains a false positive rate of at most 1.9%, thereby effectively balancing privacy preservation with auditability.

agent accountabilitycapability leakageexecution traces

This study investigates the impact of unsafe image proportions in training data on the safety of text-to-image generative models. By constructing controlled datasets and training multiple model variants with contamination rates ranging from 0% to 9.6% while holding other factors constant, the authors evaluate output safety using four independent safety classifiers, ablation studies of text encoders (including SafeCLIP), and quality metrics such as FID, CLIPScore, and ImageReward. They reveal, for the first time, a monotonic dose–response relationship between training contamination level and output unsafety. Notably, even with zero contamination, a baseline risk of 16.6% persists, which SafeCLIP reduces to 9.6% without compromising generation quality. These findings demonstrate that the text encoder itself constitutes an inherent safety risk, offering a new perspective for safer model design.

data contaminationsafety risktext-to-image models

RedactionBench

Jun 17, 2026

This study addresses the conflation of personally identifiable information recognition with privacy semantics in existing redaction benchmarks, which often neglect the critical role of context in privacy judgments. Drawing on contextual integrity theory, the authors construct RedactionBench—a human-annotated benchmark comprising 200 real-world documents across 11 domains—and introduce R-Score, a character-level evaluation metric that distinguishes semantically equivalent redactions while disregarding superficial formatting differences. The work presents the first context-aware redaction evaluation framework and systematically assesses 35 models, including named entity recognition systems, small language models, and large language models with agent-based reasoning. Human evaluations reveal low inter-annotator agreement (47.7%) on context-sensitive redaction decisions, and all evaluated models struggle significantly with this task, highlighting fundamental limitations in current approaches and underscoring the need for standardized evaluation of privacy-preserving systems.

benchmarkcontextual integritypersonally identifiable information

This work addresses the critical issue that large language models (LLMs) often reproduce security vulnerabilities present in their training data during code generation. While existing inference-time hardening techniques incur runtime overhead and cannot alter the model’s internal knowledge, this study systematically evaluates model editing as a model-level hardening mechanism for secure code generation. The authors propose SafeEdit, a novel approach that integrates task-oriented fine-tuning with edit-aware regularization to effectively mitigate the trade-off between enhanced security and functional correctness. Extensive experiments across eight prominent LLMs demonstrate that SafeEdit significantly outperforms baseline methods, achieving up to a 15.50 percentage point improvement in Pass@1 accuracy and a 7.54%–12.04% increase in security rate over CoSec, while setting new state-of-the-art performance in joint security and functionality preservation.

functional correctnesslarge language modelsmodel editing

This work addresses the vulnerability of retrieval-augmented generation (RAG) and tool-augmented large language models to malicious instructions embedded in external text, which can trigger harmful behaviors. Existing defense mechanisms suffer from poor generalization and susceptibility to optimization-based attacks. To overcome these limitations, the authors propose SONAR, a novel framework that integrates sentence-level relational graphs with natural language inference (NLI). By leveraging entailment and contradiction scores to detect malicious content and applying a connectivity-driven pruning strategy, SONAR achieves effective instruction sanitization without requiring any model retraining. Evaluated across multiple models and datasets, the method reduces attack success rates to near zero and substantially outperforms nine state-of-the-art baseline defenses.

adversarial attacksLLM agentsmalicious instructions

Hot Scholars

CX

Chaowei Xiao

University of Wisconsin - Madison/NVIDIA
Trustworthy Machine LearningAdversarial Machine LearningAI SafetyRobust AI
BA

Basel Alomair

King Abdulaziz City for Science and Technology & University of Washington
Information Security and Cryptography
EB

Elisa Bertino

Professor of Computer Science, Purdue University
Network and Computer SecurityDatabase Systems and ServicesData Privacy
JS

Jiawen Shi

Huazhong University of Science and Technology
AI Security
NZ

Neil Zhenqiang Gong

Associate Professor, Duke University
SecurityAI Security/SafetySocial Networks SecurityGenerative AI