🤖 AI Summary
This study addresses the comparative efficacy of representation engineering versus behavioral alignment for safety control and monitoring in large language models (LLMs). To this end, it proposes a unified evaluation framework that systematically contrasts Direct Preference Optimization with representation steering for safety control, alongside internal probing versus text-based monitors for safety surveillance, assessing their robustness, computational costs, and complementarity. The findings reveal that representation engineering offers distinct advantages in low-data control and cost-efficient monitoring scenarios. Crucially, this work demonstrates that representation engineering does not supplant behavioral alignment but rather complements it effectively, synergistically enhancing the overall safety of LLMs.
📝 Abstract
Reliable AI safeguards require both control mechanisms that reduce unsafe behavior and monitoring mechanisms that detect safety risks during model interactions. Established behavioral safeguards include alignment methods that optimize model outputs and text monitors that assess interaction text. Representation engineering instead reads or modifies internal model states, but the relative strengths of these approaches remain unclear because they are often evaluated under different settings. We present a matched evaluation across two tracks. For safety control, we compare DPO, a behavioral alignment method, with three representation steering methods across robustness, practicality, and granularity. DPO provides the strongest overall control and generally improves with increasing training data, although its safety can degrade after subsequent benign fine-tuning. Representation steering remains competitive primarily in low-data settings, particularly with high-quality contrastive data. For safety monitoring, we compare representation probes with fine-tuned and open-weight text monitors across full-response detection, early detection, and computational cost. Specialized text monitors achieve the strongest overall detection accuracy, while representation probes remain competitive at substantially lower marginal cost. Finally, monitor-guided interventions recover much of the safety lost by DPO after benign fine-tuning, with little additional over-refusal. Overall, representation engineering does not generally replace behavioral safeguards, but offers practical advantages under specific conditions and can provide complementary safety benefits.