π€ AI Summary
This study addresses the significant limitations of existing single-observation criteria in evaluating the stochasticity of large language model (LLM) refusal behaviors, demonstrating that a single query cannot yield accurate assessments. To overcome this, we propose a longitudinal auditing framework grounded in Bernoulli process modeling, conducting large-scale repeated prompting experiments on GPT-4.1 to quantitatively analyze the stability and decision boundaries of its refusal behavior. Our findings reveal that approximately 20% of queries reside within the decision boundary region, indicating that single observations severely underestimate the sample size required for reliable evaluation. Consequently, we establish a new standard requiring 15 to 25 repeated queries to robustly quantify model refusal behaviors, thereby providing a more rigorous methodological foundation for LLM safety assessment.
π Abstract
We present preliminary empirical evidence that single-observation queries are insufficient for evaluations of LLM refusal behaviors. Using a longitudinal auditing system, we issued identical prompts 100 times each across four dates to GPT-4.1 for two socially salient topics across 20 Wikipedia sources. Refusal outcomes were consistent with a stable Bernoulli process, yet 20\% of sources fell within a decision-boundary region where a single query is largely uninformative. Reliable quantification of refusals required between 15 and 25 repeated queries, well above the single-observation standard common in existing evaluations.