🤖 AI Summary
This study addresses the measurement inconsistency between judgment and action interfaces in LLM safety evaluation by proposing SameFact, the first matched counterfactual benchmark. Comprising 300 safe/unsafe paired instances, SameFact systematically evaluates six LLMs across three response interfaces—judgment, checkpoint, and open-choice—by holding non-target variables constant while altering a single safety-relevant fact. The findings reveal that the response interface itself constitutes a core element of measurement, demonstrating that judgment and action interfaces are not interchangeable. Experimental results indicate that sensitivity and ranking consistency degrade significantly under the open-choice interface, whereas the checkpoint protocol improves measurement sensitivity by 8.4 to 29.3 percentage points.
📝 Abstract
Safety evaluations often ask whether a model recognizes that an action is unsafe, whereas agent evaluations ask what the model chooses to do. Using safety judgments as evidence about action selection therefore raises a measurement question: does the influence of the same safety-relevant fact persist across response interfaces? We introduce SameFact, a matched-counterfactual benchmark that tests this question directly. SameFact contains 300 safe/unsafe pairs that hold the task, prior observations, candidate action, identifiers, and non-target facts fixed while changing a single state-grounded safety fact. Across six LLM backbones, we measure the effect of this matched intervention through three interfaces at the same candidate-action boundary: explicit safety judgment, checkpoint candidate admission, and open first-action selection. All six backbones show lower aggregate sensitivity under open first-action selection than under judgment, but the change is not a uniform attenuation: across 24 model-factor cells, Spearman agreement falls from 0.817 between judgment and checkpoint admission to 0.470 between judgment and open first-action selection, while pairwise ordering disagreement rises from 18.5% to 32.6%. A follow-up 2x2 first-response experiment shows that a checkpoint-style protocol increases measured sensitivity in all six backbones by 8.4-29.3 percentage points, whereas action-space effects and their interactions with protocol vary in magnitude and direction across models. These results show that the response interface is part of the measured quantity: judgment and action interfaces share safety signal, but do not provide interchangeable measurements of how safety-relevant facts shape model responses.