EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures

📅 2026-06-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses a critical “evaluation–safety gap” (EvalSafetyGap) in large language models (LLMs), wherein apparent performance gains do not necessarily reflect genuine safety capabilities. Through a systematic literature review, gray literature analysis, and a multidimensional audit of ten models across eight evidence streams, the work proposes the EvalSafetyGap hypothesis and introduces two novel constructs—“instability decomposition” and the “alignment trilemma”—to establish a unified terminology and evidence map supporting dynamic evaluation and auditable alignment. Empirical findings reveal no significant correlation between model capability and adversarial robustness (r = 0.232, p = 0.520). Moreover, safety differences between open- and closed-source models stem primarily from governance transparency rather than behavioral robustness, with results highly sensitive to model categorization and evaluation protocols.
📝 Abstract
LLM evaluation and AI safety face a shared measurement problem: benchmark scores, reward-model signals, and reported safety metrics can improve while the latent properties they are meant to represent remain difficult to verify. This paper combines a hybrid survey - a systematic search paired with narrative synthesis and separately tracked grey evidence - with a conceptual framework and a structured ten-model audit. The synthesis spans eight evidence streams: benchmark validity, dynamic evaluation, LLM-as-judge reliability, safety evaluation, jailbreak/refusal robustness, reward hacking, mechanistic interpretability, and governance/auditability, covering 2018-2026 evaluation-safety measurement work. We introduce EvalSafetyGap as an organizing hypothesis for comparing evaluation-side and alignment-side proxy failures under optimization pressure, using Goodhart's Law together with two constructs we develop here - an Instability Decomposition and an Alignment Trilemma - as tools for generating testable comparisons. The audit shows how conclusions shift when capability, behavioral safety, and governance are measured separately. In this sample (n = 10), the association between capability and sustained adversarial robustness is statistically indeterminate using the displayed Table 3 inputs (Pearson r = +0.232, p = 0.520), and the apparent open-closed safety gap is modest, driven mainly by governance and disclosure rather than behavioral robustness, and sensitive to how a single borderline model is classified; attempt-budget results are protocol dependent. Because the public evidence uses heterogeneous protocols, the audit is diagnostic rather than rank-generating. The contribution is a shared vocabulary and evidence map to support dynamic evaluation, transparent source reporting, multi-attempt safety measurement, and auditable alignment practice.
Problem

Research questions and friction points this paper is trying to address.

LLM evaluation
AI safety
measurement gap
benchmark validity
alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

EvalSafetyGap
Alignment Trilemma
Instability Decomposition
dynamic evaluation
auditable alignment
B
Buğra Alperen Uluırmak
Erciyes University
R
Rifat Kurban
Abdullah Gül University