Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety

📅 2026-03-08
📈 Citations: 0
Influential: 0
📄 PDF

career value

183K/year
🤖 AI Summary
This study addresses the limitations of current safety evaluations, which predominantly rely on isolated multiple-choice setups and overlook the real-world impact of agent scaffolding on model safety. Through a large-scale controlled experiment (N = 62,808), we systematically assess four scaffolding architectures—including Map-Reduce—across evaluation formats (open-ended vs. multiple-choice) on state-of-the-art language models. We find that evaluation format exerts a far stronger influence on safety scores than scaffolding effects themselves. Critically, strong model–scaffolding interactions lead to complete reversals in safety rankings across benchmarks (G = 0.000). Employing preregistration, evaluator blinding, TOST equivalence testing, and generalizability analyses, we demonstrate that safety must be evaluated for each specific model–deployment configuration: Map-Reduce significantly reduces safety (NNH = 14), whereas other architectures remain equivalent within ±2 percentage points.

Technology Category

Application Category

📝 Abstract
Safety benchmarks evaluate language models in isolation, typically using multiple-choice format; production deployments wrap these models in agentic scaffolds that restructure inputs through reasoning traces, critic agents, and delegation pipelines. We report one of the largest controlled studies of scaffold effects on safety (N = 62,808; six frontier models, four deployment configurations), combining pre-registration, assessor blinding, equivalence testing, and specification curve analysis. Map-reduce scaffolding degrades measured safety (NNH = 14), yet two of three scaffold architectures preserve safety within practically meaningful margins. Investigating the map-reduce degradation revealed a deeper measurement problem: switching from multiple-choice to open-ended format on identical items shifts safety scores by 5-20 percentage points, larger than any scaffold effect. Within-format scaffold comparisons are consistent with practical equivalence under our pre-registered +/-2 pp TOST margin, isolating evaluation format rather than scaffold architecture as the operative variable. Model x scaffold interactions span 35 pp in opposing directions (one model degrades by -16.8 pp on sycophancy under map-reduce while another improves by +18.8 pp on the same benchmark), ruling out universal claims about scaffold safety. A generalisability analysis yields G = 0.000: model safety rankings reverse so completely across benchmarks that no composite safety index achieves non-zero reliability, making per-model, per-configuration testing a necessary minimum standard. We release all code, data, and prompts as ScaffoldSafety.
Problem

Research questions and friction points this paper is trying to address.

safety evaluation
language models
scaffolding
evaluation format
benchmarking
Innovation

Methods, ideas, or system contributions that make the work stand out.

scaffold safety
evaluation format
agentic scaffolding
equivalence testing
model safety benchmarking
🔎 Similar Papers