🤖 AI Summary
This study addresses the significant disconnect between safety judgment and behavioral preference in large language model (LLM) agents, which can recognize hazardous actions yet still execute them. Through intervention experiments, causal subspace analysis, and cross-model comparisons, we investigate the underlying mechanisms of this phenomenon, revealing a fundamental distinction between information availability and causal control. Our findings demonstrate that strong causal control within the judgment pathway fails to translate into effective constraints on the action pathway, as their representational subspaces only partially overlap. This work indicates that merely enhancing self-critique capabilities is insufficient for ensuring agent safety, thereby providing crucial theoretical foundations for developing reliable alignment mechanisms.
📝 Abstract
Large language model agents can correctly judge that an action should be blocked while still preferring to take it. We ask why this judgment-action disconnect arises, and whether explicit safety judgment causally governs subsequent action preference. Across three open-weight language models, safety-predictive information remains recoverable from action states, arguing against a simple information-loss account. Instead, the disconnect is better explained by weak coupling between judgment- and action-side causal control: interventions that reliably shift explicit safety judgments toward BLOCK produce much smaller changes in action preference than action-native interventions. This asymmetry persists within a shared judgment-to-action trajectory, where strong upstream control of judgment does not translate into comparably strong downstream control of action preference. Beyond individual intervention directions, judgment- and action-control subspaces overlap only partially, while effective action control remains available in directions orthogonal to the judgment-control subspace. Together, these results distinguish information availability from causal control: an LLM agent can retain the information needed to recognize an action as unsafe without the variables supporting that judgment reliably governing its action preference. For agent safety, this suggests that improving safety recognition or self-critique alone may be insufficient unless safety-relevant computations are also causally coupled to action selection.