inference-time concept erasure

Designs and implements operators and procedures that remove, neutralize, or project out specific concept-related signals from a model’s internal activations at inference time without retraining; this includes closed-form, training-free projection methods (e.g., CARE operators) that modify activations and accompanying evaluation workflows to measure model behavior after concept erasure.

inference-timeconcepterasure

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.62
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

CURE: Concept Unlearning via Orthogonal Representation Editing in Diffusion Models

May 19, 2025
SD
Shristi Das Biswas
🏛️ Purdue University

To address the challenge of precisely suppressing unsafe, copyrighted, or privacy-invasive concepts in pre-trained diffusion models, this paper proposes a training-free concept unloading framework. Our method directly identifies the token embedding subspace associated with the target concept in the weight space, analyzes its spectral structure via Singular Value Decomposition (SVD), and designs a “Spectral Eraser” that performs closed-form orthogonal projection to selectively suppress harmful concepts. The entire process requires no fine-tuning, supervision, or iterative optimization, completing editing in under two seconds. Experiments on artistic style, object, identity, and explicit content removal demonstrate substantial improvements over baselines, achieving high generation fidelity, minimal capability degradation, and strong robustness against red-teaming jailbreak attacks. The core contribution lies in the first integration of spectral analysis with closed-form weight-space editing—enabling fast, interpretable, and high-precision targeted forgetting.

Balancing toxicity filtering with unrelated concept preservationImproving efficiency and specificity of concept removalPreventing unsafe or copyrighted content in diffusion models

A Granular Study of Safety Pretraining under Model Abliteration

Oct 03, 2025
SA
Shashank Agnihotri
🏛️ University of Mannheim | Indian Institute of Science | Carnegie Mellon University | Max-Planck-Institute for Informatics

This study investigates the robustness of mainstream safety interventions—such as refusal training and meta-label training—in open-weight large language models under lightweight activation-based editing (i.e., model pruning) at inference time. We systematically evaluate how well multiple safety-pretrained checkpoints retain their refusal behavior post-pruning, introducing a fine-grained checkpoint analysis framework that integrates self-referential refusal detection, multi-judge classification, and human-annotated validation to establish a joint protocol for inference-time editing and safety assessment. Results show that certain safety mechanisms degrade significantly under pruning, with refusal-sensitive directions particularly vulnerable to removal; judge selection substantially impacts evaluation outcomes; and data-driven safety components exhibit pronounced inter-checkpoint capability variance. This work provides the first quantitative evidence of how pruning undermines safety alignment, empirically delineating the safety boundaries of inference-time model editing.

Assessing refusal behavior survival under lightweight projection techniquesCharacterizing checkpoint-level safety component resilience to inference-time modificationsEvaluating safety training robustness against model activation edits

This work addresses the challenge of efficiently and safely removing specific concepts from generative models without affecting unrelated content. It proposes a training-free, closed-form linear transformation framework that achieves concept erasure through a two-step analytical projection: first computing a proxy projection of the target concept, then applying a constrained transformation within its left null space. As the first deterministic, geometrically interpretable, and non-iterative method for concept editing, it accomplishes erasure in just seconds on Stable Diffusion variants and FLUX models. The approach matches or exceeds state-of-the-art performance while significantly improving computational efficiency and better preserving the integrity of non-target concepts.

concept erasureethical risksgenerative models

This work addresses the issue that existing training-free methods for erasing specific concepts from diffusion models often inadvertently remove semantically related non-target content. To mitigate this, the authors propose CARE, a closed-form concept erasure operator that constructs a perceptually preserved subspace in the cross-attention value space, guided by anchor representations of retained concepts. The target concept direction is then replaced with its projection onto this subspace, enabling precise erasure while preserving shared visual structures. CARE incorporates an adjustable shrinkage parameter to balance erasure efficacy and semantic retention, and it offers theoretical guarantees of minimal perturbation. Experiments demonstrate that CARE significantly outperforms current state-of-the-art methods across instance-, style-, and celebrity-level concept erasure tasks, while effectively safeguarding unrelated semantic information.

collateral damageconcept erasurecross-attention

Intrinsic Evaluation of Unlearning Using Parametric Knowledge Traces

Jun 17, 2024
YH
Yihuai Hong
🏛️ South China University of Technology | University of Toronto | Bar-Ilan University | International Digital Economy Academy (IDEA) | Tel Aviv University

Existing LLM unlearning evaluations rely solely on behavioral testing, overlooking residual knowledge at the parameter level—enabling adversarial recovery of supposedly deleted information. This work proposes the first intrinsic unlearning evaluation framework grounded in parameter-level knowledge traces: it localizes “concept vectors” via vocabulary projection, constructs the open-source benchmark ConceptVectors, and systematically models knowledge traces in Llama-2 and Phi-3. We find that mainstream unlearning methods merely suppress—not erase—concept vectors; direct ablation of these vectors fully removes associated knowledge and drastically reduces adversarial recovery success rates. Our study establishes a new paradigm for parameter-level unlearning assessment, exposes fundamental limitations of behavioral evaluation, and advances unlearning research from black-box testing toward mechanistic interpretability. Code and the ConceptVectors benchmark are publicly released.

Detecting residual knowledge in model parameters post-unlearningEvaluating unlearning methods beyond behavioral testsLocalizing concept vectors to assess knowledge removal effectiveness

Latest Papers

What's happening recently
View more

This work addresses the pressing need for efficient, reversible unlearning methods that avoid retraining large language models, which inherently memorize training data and thereby pose privacy, copyright, and security risks. The authors propose a training- and gradient-free inference-time unlearning mechanism that adaptively gates activations to apply norm-preserving rotational transformations in the residual stream, precisely removing the influence of specified data without altering model weights. This approach enables, for the first time, dynamic and localized activation steering, circumventing the side effects of global interventions and supporting continual unlearning even in quantized models. Evaluated on TOFU and MUSE benchmarks across three model scales, the method consistently outperforms twelve gradient-based baselines, effectively suppressing memorization while preserving model utility and maintaining robustness under quantization.

activation steeringinference-time interventionlarge language models

This study investigates whether large language models exhibit operation-level causal transfer structures across finite isomorphic symbolic domains. Focusing on Qwen2.5, we propose a "predefined construction-confirmation" testing paradigm to rigorously decouple validation phases, employing input-specific activation interventions alongside dual-framework causal analysis using PyVene and NNsight. Our results confirm the existence of operation-level causal transfer along specific computational paths, demonstrating statistical significance and numerical reproducibility across both frameworks. By validating these internal mechanisms, this work provides rigorous empirical evidence and methodological support for understanding cross-domain causal transfer within large language models, advancing the mechanistic interpretability of their symbolic reasoning capabilities.

Construction-confirmation testIsomorphic state spacesOperation-level causal transfer

Frozen small code models often generate programs that appear plausible yet are incorrect, necessitating post-hoc correction mechanisms that do not require fine-tuning. This work systematically evaluates 26 semantic post-processing operators and finds they generally fail to surpass the Best-of-N baseline, identifying three key bottlenecks: the coverage wall, capability scissors, and consensus trap. To address these limitations, the paper introduces two harmless and effective non-semantic methods: Manifestation-layer Recovery (M1) and Adaptive Cutoff Early-stopping (ACE). M1 significantly improves performance on 12 HumanEval+ tasks for DeepSeek-Coder-1.3B (p = 2.4e-4), while ACE reduces computational cost by approximately 19%. Both methods achieve zero harm and demonstrate consistent gains across multiple models and benchmarks.

code generation accuracyfrozen small code modelsmodel output reliability

This study investigates whether the self-repair capability of frozen small code models in non-retrainable settings stems from repeated exposure to failed code or relies on external executable falsification feedback. To address this, we introduce a falsifiable methodology comprising feedback decomposition, content-controlled placebo design, matched-generation-budget control experiments, and executable auditing. We conduct large-scale evaluations on HumanEval+ and MBPP+ benchmarks using frozen models ranging from 0.5B to 1.5B parameters. Results show that blind resampling solves 18 more tasks than naive retrying; significant repair efficacy occurs only when feedback includes executable counterexamples, whereas pure instructions or content-irrelevant placebos yield no measurable improvement. These findings demonstrate that effective self-repair depends critically on external falsifying information rather than mere self-restatement.

falsificationfeedback decompositionfrozen code models

Hot Scholars

PG

Philipp G. Haselwarter

Assistant Professor, Aarhus University
Programming LanguagesLogicCryptographyType Theory
PL

Ping Liu

Assistant Professor, Krannert School of Management, Purdue University
Contract theoryGame theoryMacro financeReal Options
HT

Hao Tang

University of Edinburgh
Speech and Language ProcessingSpeech Recognition