🤖 AI Summary
This study addresses the persistent irreproducibility in LLM-as-judge safety evaluations, even when using greedy decoding (temperature = 0). Contrary to common assumptions, the authors demonstrate that temperature = 0 does not fully eliminate stochasticity in LLM scoring, particularly for borderline cases. Through 690 experiments across multiple models, APIs, and sampling configurations within the open-source aisev evaluation framework, they reveal that default temperature settings can induce judgment fluctuations in up to 50% of individual items, and even deterministic decoding fails to ensure reproducibility for one to two borderline cases per evaluation. To address this, the paper proposes incorporating scoring disagreement as a core health metric in evaluation frameworks, advocating for more robust and reliable practices in LLM safety assessment.
📝 Abstract
LLM-as-judge ("grader") components are now standard in evaluation harnesses, including safety evaluations where a pass/fail verdict may gate downstream deployment decisions. A widespread assumption is that setting the grader's sampling temperature to 0 makes grading deterministic. We test this assumption against a real safety-evaluation codebase (Japan AISI's open-source aisev) and show it fails on two levels. First, the harness invokes its grader without setting temperature or seed; the underlying provider silently applies its default of 1.0, so items near the decision boundary flip pass/fail across identical runs (per-item disagreement up to ~50% over 20 runs). Second, pinning temperature=0 reduces but does not eliminate flips: across 690 API calls spanning two providers, three model tiers, and five sampling configurations, 1-2 of 7 borderline items remain non-reproducible even under forced greedy decoding (top_k=1). Claude Opus 4.7/4.8 has since deprecated temperature entirely, rendering the primary mitigation inapplicable to newer model generations. These findings expose a structural gap: evaluation harnesses that report single-run verdicts without variance or grader-disagreement metrics can present noise as a safety property. We release a reproduction harness (690 calls, 7 conditions) and recommend that harnesses treat grader disagreement as a first-class health metric alongside the scores themselves.