Don't Inoculate Everything: Stratified Inoculation Prompting Narrows Backdoor Triggers and Preserves Desired Traits
This study addresses the challenges of undesirable behavior generalization and suppressed desired learning during supervised fine-tuning. To mitigate these issues, it proposes a hierarchical inoculation prompting framework that leverages a small set of clean samples to reinforce desired behaviors across diverse contexts while isolating undesirable ones. Furthermore, backdoor dilution and password-locking mechanisms are introduced to enable selective generalization, complemented by hierarchical sampling, diversified non-triggering prompts, and adversarial training strategies to enhance model robustness. This approach substantially suppresses the expression of undesirable behaviors while effectively preserving desired characteristics, thereby significantly reducing the rate of emergent misalignment. Ultimately, this work establishes a novel paradigm for safe and controllable training during the alignment phase.