๐ค AI Summary
This study addresses the limitation that aggregate scores in existing AI safety benchmarks obscure differences in model properties, making it difficult to verify whether โharmful refusalโ constitutes a single measurable construct. To systematically examine this construct validity, we introduce a psychometric framework for auditing the HELM safety benchmark, employing multidimensional item response theory (MIRT) modeling, differential item functioning (DIF) analysis, and saturation testing. Our findings reveal that most datasets have already reached saturation, confirming that HarmBench fails to measure a unified safety attribute and exposing fundamental flaws in conventional aggregate scoring. Consequently, this work argues that safety evaluations must undergo rigorous attribute validation before being used for model comparison, thereby establishing a new paradigm for benchmark design.
๐ Abstract
Safety benchmarks typically report one overall score for a suite of datasets, each of which may target one or more safety-related attributes, so models with similar overall scores can have very different attribute profiles. Comparing models is more tractable at the level of individual attributes, yet it is often unclear whether even a single dataset's scores isolate any single attribute. One plausible candidate for such an attribute is harmful refusal, a model's tendency to refuse dangerous or policy-violating prompts. We examine whether it constitutes a single, measurable attribute in HELM Safety. Using a construct validity framework that stipulates that an attribute must exist before a test can measure it, we start with HELM Safety's four datasets that might plausibly target harmful refusal, but find that three are saturated. We subject the remaining dataset, HarmBench, to two psychometric tests to determine if a single attribute like harmful refusal could stand behind its score. First, multidimensional item response theory modeling strongly suggests that HarmBench does not measure a singular attribute. Second, a differential item functioning analysis finds items where models from different developers with the same refusal ability score differently. These flags largely disappear under scope-specific matching, a pattern consistent with aggregation effects but not sufficient to rule out domain-specific developer differences. Zooming out, HarmBench collapses distinct harm behaviors into one score, and the overall HELM safety aggregate further collapses HarmBench and scores from other datasets into a single top-line number. Any safety score that averages over datasets and items can hide saturation and conflate behaviors this way. We argue that a score should earn its single-attribute reading before models are compared with it.