Score
Designs and implements operators and procedures that remove, neutralize, or project out specific concept-related signals from a model’s internal activations at inference time without retraining; this includes closed-form, training-free projection methods (e.g., CARE operators) that modify activations and accompanying evaluation workflows to measure model behavior after concept erasure.
Text-to-image (T2I) diffusion models frequently generate sensitive, harmful, or copyright-protected content, necessitating controllable concept suppression mechanisms. To address this, we systematically survey existing concept removal methods and introduce the first three-dimensional taxonomy—spanning intervention level, optimization architecture, and semantic scope—that exposes fundamental trade-offs among specificity, generalization, and efficiency. We propose an evaluation gap analysis framework and establish the first unified taxonomy and benchmark specifically designed for ethical alignment. By integrating gradient-driven editing, latent-space intervention, and concept disentanglement techniques, we empirically validate fine-grained semantic masking while preserving generation fidelity. Our contributions provide a reproducible, rigorously evaluable methodological foundation for responsible generative AI.
To address the challenge of precisely suppressing unsafe, copyrighted, or privacy-invasive concepts in pre-trained diffusion models, this paper proposes a training-free concept unloading framework. Our method directly identifies the token embedding subspace associated with the target concept in the weight space, analyzes its spectral structure via Singular Value Decomposition (SVD), and designs a “Spectral Eraser” that performs closed-form orthogonal projection to selectively suppress harmful concepts. The entire process requires no fine-tuning, supervision, or iterative optimization, completing editing in under two seconds. Experiments on artistic style, object, identity, and explicit content removal demonstrate substantial improvements over baselines, achieving high generation fidelity, minimal capability degradation, and strong robustness against red-teaming jailbreak attacks. The core contribution lies in the first integration of spectral analysis with closed-form weight-space editing—enabling fast, interpretable, and high-precision targeted forgetting.
This study investigates the robustness of mainstream safety interventions—such as refusal training and meta-label training—in open-weight large language models under lightweight activation-based editing (i.e., model pruning) at inference time. We systematically evaluate how well multiple safety-pretrained checkpoints retain their refusal behavior post-pruning, introducing a fine-grained checkpoint analysis framework that integrates self-referential refusal detection, multi-judge classification, and human-annotated validation to establish a joint protocol for inference-time editing and safety assessment. Results show that certain safety mechanisms degrade significantly under pruning, with refusal-sensitive directions particularly vulnerable to removal; judge selection substantially impacts evaluation outcomes; and data-driven safety components exhibit pronounced inter-checkpoint capability variance. This work provides the first quantitative evidence of how pruning undermines safety alignment, empirically delineating the safety boundaries of inference-time model editing.
This work addresses the challenge of efficiently and safely removing specific concepts from generative models without affecting unrelated content. It proposes a training-free, closed-form linear transformation framework that achieves concept erasure through a two-step analytical projection: first computing a proxy projection of the target concept, then applying a constrained transformation within its left null space. As the first deterministic, geometrically interpretable, and non-iterative method for concept editing, it accomplishes erasure in just seconds on Stable Diffusion variants and FLUX models. The approach matches or exceeds state-of-the-art performance while significantly improving computational efficiency and better preserving the integrity of non-target concepts.
This work addresses the issue that existing training-free methods for erasing specific concepts from diffusion models often inadvertently remove semantically related non-target content. To mitigate this, the authors propose CARE, a closed-form concept erasure operator that constructs a perceptually preserved subspace in the cross-attention value space, guided by anchor representations of retained concepts. The target concept direction is then replaced with its projection onto this subspace, enabling precise erasure while preserving shared visual structures. CARE incorporates an adjustable shrinkage parameter to balance erasure efficacy and semantic retention, and it offers theoretical guarantees of minimal perturbation. Experiments demonstrate that CARE significantly outperforms current state-of-the-art methods across instance-, style-, and celebrity-level concept erasure tasks, while effectively safeguarding unrelated semantic information.
Existing LLM unlearning evaluations rely solely on behavioral testing, overlooking residual knowledge at the parameter level—enabling adversarial recovery of supposedly deleted information. This work proposes the first intrinsic unlearning evaluation framework grounded in parameter-level knowledge traces: it localizes “concept vectors” via vocabulary projection, constructs the open-source benchmark ConceptVectors, and systematically models knowledge traces in Llama-2 and Phi-3. We find that mainstream unlearning methods merely suppress—not erase—concept vectors; direct ablation of these vectors fully removes associated knowledge and drastically reduces adversarial recovery success rates. Our study establishes a new paradigm for parameter-level unlearning assessment, exposes fundamental limitations of behavioral evaluation, and advances unlearning research from black-box testing toward mechanistic interpretability. Code and the ConceptVectors benchmark are publicly released.
This work addresses the pressing need for efficient, reversible unlearning methods that avoid retraining large language models, which inherently memorize training data and thereby pose privacy, copyright, and security risks. The authors propose a training- and gradient-free inference-time unlearning mechanism that adaptively gates activations to apply norm-preserving rotational transformations in the residual stream, precisely removing the influence of specified data without altering model weights. This approach enables, for the first time, dynamic and localized activation steering, circumventing the side effects of global interventions and supporting continual unlearning even in quantized models. Evaluated on TOFU and MUSE benchmarks across three model scales, the method consistently outperforms twelve gradient-based baselines, effectively suppressing memorization while preserving model utility and maintaining robustness under quantization.
This study investigates whether large language models exhibit operation-level causal transfer structures across finite isomorphic symbolic domains. Focusing on Qwen2.5, we propose a "predefined construction-confirmation" testing paradigm to rigorously decouple validation phases, employing input-specific activation interventions alongside dual-framework causal analysis using PyVene and NNsight. Our results confirm the existence of operation-level causal transfer along specific computational paths, demonstrating statistical significance and numerical reproducibility across both frameworks. By validating these internal mechanisms, this work provides rigorous empirical evidence and methodological support for understanding cross-domain causal transfer within large language models, advancing the mechanistic interpretability of their symbolic reasoning capabilities.
论文解决了大语言模型后训练中经验重用的问题,提出了一种名为BCIT的方法来有条件地转移经验,避免无效或有害的更新,从而提高模型质量。
Frozen small code models often generate programs that appear plausible yet are incorrect, necessitating post-hoc correction mechanisms that do not require fine-tuning. This work systematically evaluates 26 semantic post-processing operators and finds they generally fail to surpass the Best-of-N baseline, identifying three key bottlenecks: the coverage wall, capability scissors, and consensus trap. To address these limitations, the paper introduces two harmless and effective non-semantic methods: Manifestation-layer Recovery (M1) and Adaptive Cutoff Early-stopping (ACE). M1 significantly improves performance on 12 HumanEval+ tasks for DeepSeek-Coder-1.3B (p = 2.4e-4), while ACE reduces computational cost by approximately 19%. Both methods achieve zero harm and demonstrate consistent gains across multiple models and benchmarks.
This study investigates whether the self-repair capability of frozen small code models in non-retrainable settings stems from repeated exposure to failed code or relies on external executable falsification feedback. To address this, we introduce a falsifiable methodology comprising feedback decomposition, content-controlled placebo design, matched-generation-budget control experiments, and executable auditing. We conduct large-scale evaluations on HumanEval+ and MBPP+ benchmarks using frozen models ranging from 0.5B to 1.5B parameters. Results show that blind resampling solves 18 more tasks than naive retrying; significant repair efficacy occurs only when feedback includes executable counterexamples, whereas pure instructions or content-irrelevant placebos yield no measurable improvement. These findings demonstrate that effective self-repair depends critically on external falsifying information rather than mere self-restatement.