🤖 AI Summary
This study addresses the challenges of evidence accumulation and verification budget determination in third-party validation caused by the stochastic outputs of probabilistic AI models. To overcome these issues, this work quantifies the statistical separability within challenge responses and establishes a theoretical linkage between single-instance behavior and aggregate detection performance. By leveraging open-domain testing with large language models and area under the curve (AUC) theoretical modeling, it derives an explicit estimation method for determining the minimum verification budget required to achieve a target AUC. The proposed approach effectively separates matching from non-matching distributions, demonstrating strong agreement between theoretical predictions and empirical AUC values. Ultimately, this research provides a precise budget planning framework that facilitates the reliable verification of probabilistic AI models.
📝 Abstract
Third-party challenge-response identity verification (TP-CRIV) enables an independent verifier to assess whether a claimant possesses a model identical to a remotely deployed model without directly accessing the reference model. However, for probabilistic AI models, repeated executions of the same query may produce different outputs and therefore different verification observations. This raises the question of how such stochastic evidence should be accumulated and how much evidence is required for reliable verification.
In this work, we characterize statistical separability in TP-CRIV of probabilistic AI models. Specifically, we relate challenge-wise behavior of matching and non-matching provers to verification-level separability. The characterization explicitly describes how the numbers of independent challenges and repeated responses affect detection performance and enables the verification budget required for a target AUC to be estimated. We instantiate the proposed characterization for LLMs using open-ended challenges. The experiments demonstrate matching-non-matching separation, close agreement between theoretical and empirical AUCs, and consistent estimates of the minimum verification budgets. These results provide a statistical basis for relating probabilistic model behavior to verification-level separability and the evidence required for third-party verification.