Institution profile

Center on Long-Term Risk

Academic institutioneurope · gb
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Don't Inoculate Everything: Stratified Inoculation Prompting Narrows Backdoor Triggers and Preserves Desired Traits

Sep 28, 2026

This study addresses the challenges of undesirable behavior generalization and suppressed desired learning during supervised fine-tuning. To mitigate these issues, it proposes a hierarchical inoculation prompting framework that leverages a small set of clean samples to reinforce desired behaviors across diverse contexts while isolating undesirable ones. Furthermore, backdoor dilution and password-locking mechanisms are introduced to enable selective generalization, complemented by hierarchical sampling, diversified non-triggering prompts, and adversarial training strategies to enhance model robustness. This approach substantially suppresses the expression of undesirable behaviors while effectively preserving desired characteristics, thereby significantly reducing the rate of emergent misalignment. Ultimately, this work establishes a novel paradigm for safe and controllable training during the alignment phase.

0 citationsRead paper

Inoculation Adapters: Improved Selective Generalization of Capabilities with Fewer Surprising Backdoors

Jun 29, 2026

This work addresses the challenge of suppressing undesirable behaviors—such as sudden alignment failures—learned during model training while preserving desired capabilities and avoiding unintended backdoors. The authors propose the Inoculation Adapter (IA) method, which first trains a LoRA adapter specialized in capturing undesirable behaviors, then freezes this adapter to guide the training of the main task adapter. Only the main adapter is deployed, thereby reducing the optimization pressure that leads the model to acquire undesirable capabilities. Unlike prompt-based inoculation, IA effectively mitigates behaviors that are difficult to elicit via prompting and substantially diminishes the risk of accidental backdoors. Experiments across six model families demonstrate that IA achieves more selective capability suppression while enhancing both safety and general applicability.

0 citationsRead paper
Recent publications

Latest Papers

Don't Inoculate Everything: Stratified Inoculation Prompting Narrows Backdoor Triggers and Preserves Desired Traits

Sep 28, 2026

This study addresses the challenges of undesirable behavior generalization and suppressed desired learning during supervised fine-tuning. To mitigate these issues, it proposes a hierarchical inoculation prompting framework that leverages a small set of clean samples to reinforce desired behaviors across diverse contexts while isolating undesirable ones. Furthermore, backdoor dilution and password-locking mechanisms are introduced to enable selective generalization, complemented by hierarchical sampling, diversified non-triggering prompts, and adversarial training strategies to enhance model robustness. This approach substantially suppresses the expression of undesirable behaviors while effectively preserving desired characteristics, thereby significantly reducing the rate of emergent misalignment. Ultimately, this work establishes a novel paradigm for safe and controllable training during the alignment phase.

0 citationsRead paper

Inoculation Adapters: Improved Selective Generalization of Capabilities with Fewer Surprising Backdoors

Jun 29, 2026

This work addresses the challenge of suppressing undesirable behaviors—such as sudden alignment failures—learned during model training while preserving desired capabilities and avoiding unintended backdoors. The authors propose the Inoculation Adapter (IA) method, which first trains a LoRA adapter specialized in capturing undesirable behaviors, then freezes this adapter to guide the training of the main task adapter. Only the main adapter is deployed, thereby reducing the optimization pressure that leads the model to acquire undesirable capabilities. Unlike prompt-based inoculation, IA effectively mitigates behaviors that are difficult to elicit via prompting and substantially diminishes the risk of accidental backdoors. Experiments across six model families demonstrate that IA achieves more selective capability suppression while enhancing both safety and general applicability.

0 citationsRead paper