π€ AI Summary
This study addresses the degradation of general capabilities and output naturalness caused by single-objective optimization in model organism training, which compromises the accuracy of interpretability evaluations. To overcome this limitation, we propose a multi-objective training framework based on model merging. This approach introduces a tri-objective validation metric and integrates Direct Preference Optimization (DPO) with supervised fine-tuning comparisons to enhance behavioral fidelity while effectively preserving the base modelβs general capabilities, dialogue quality, and output naturalness. Experimental results demonstrate that the proposed validation metric accurately predicts behavioral recovery outcomes, and the new training paradigm significantly outperforms conventional approaches. Ultimately, this work provides a more reliable model foundation for white-box interpretability research.
π Abstract
Model organisms of alignment-relevant behaviors (e.g., backdoors, sycophancy, spurious correlations) have emerged as a key tool for evaluating whitebox interpretability techniques. We argue that the prevailing practice of training model organisms to a single objective of installing the target behavior is insufficient and propose validating model organisms with respect to three objectives with associated metrics: target-behavior installation, general-capability preservation (i.e., parametric knowledge, chat quality), and output naturalness (i.e., CoT and activations). We re-visit two publicly released organism suites using this validation framework and show that (1) chat quality and CoT naturalness degrade substantially across training recipes, and (2) validation metrics predict how well interpretability methods recover the installed behavior, e.g., a logit lens readout covaries with an organism's general capabilities. We introduce a multi-objective training approach based on model merging to train more realistic model organisms. Finally, on a new suite of model organisms targeting demographic biases in clinical reasoning, we compare training recipes and find that DPO training stays closer to the base model than supervised finetuning, and the proposed model optimization approach better preserves capabilities and naturalness. Auditing this suite with an investigator agent, we again observe validation metrics tracking bias recovery. In sum, training methods shape the interpretability conclusions an organism supports, and we argue that one should consider multiple objectives to draw generalizable conclusions about interpretability methods using (realistic) model organisms.