Item Response Theory for AI Safety

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing AI safety evaluation benchmarks suffer from redundancy, high inter-correlation, and potential sandbagging—where models deliberately underperform—rendering aggregated scores difficult to interpret and trust. This work presents the first large-scale application of Item Response Theory (IRT) to safety assessment of large language models (LLMs), integrating factor analysis and adaptive testing to distill three interpretable latent traits from multiple benchmarks: refusal strictness, truthfulness, and contextual harm. The proposed IRT-based framework achieves 97–99% fidelity in replicating individual benchmark outcomes using only about ten adaptively selected items, substantially improving evaluation efficiency. Furthermore, it effectively detects sandbagging behaviors, demonstrating IRT’s reliability, scalability, and auditability in LLM safety evaluation.
📝 Abstract
Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores are hard to trust and interpret, because benchmarks duplicate one another, correlate heavily, and models may sandbag when they detect evaluation. To address these issues, we draw on Item Response Theory (IRT), a statistical toolkit for measuring these latents from performance on items with inferred psychometric properties. We fit IRT models to eight safety benchmarks across 192 language models, the largest psychometric analysis of LLM safety evaluations to date, and contribute three results. First, we find that three interpretable factors of refusal strictness, truthfulness, and contextual harm explain most of the variance between models across benchmarks. Second, psychometrically selected items recover full benchmark scores with lower error than random subsets of the same size, and roughly ten adaptively chosen items suffice for several individual benchmarks, cutting evaluation cost by 97-99%. Third, IRT supports audits of individual models, showing that it can be used to detect naive sandbagging and changes of model behind APIs. Overall, we show IRT is a ready-made toolkit for reading, reducing, and auditing safety benchmarks, which we recommend frontier labs and evaluators adopt.
Problem

Research questions and friction points this paper is trying to address.

AI safety
safety benchmarks
language models
evaluation reliability
sandbagging
Innovation

Methods, ideas, or system contributions that make the work stand out.

Item Response Theory
AI Safety
Psychometric Analysis
Adaptive Testing
Model Auditing
🔎 Similar Papers