Characterizing Statistical Separability in TP-CRIV for Probabilistic AI Models

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of evidence accumulation and verification budget determination in third-party validation caused by the stochastic outputs of probabilistic AI models. To overcome these issues, this work quantifies the statistical separability within challenge responses and establishes a theoretical linkage between single-instance behavior and aggregate detection performance. By leveraging open-domain testing with large language models and area under the curve (AUC) theoretical modeling, it derives an explicit estimation method for determining the minimum verification budget required to achieve a target AUC. The proposed approach effectively separates matching from non-matching distributions, demonstrating strong agreement between theoretical predictions and empirical AUC values. Ultimately, this research provides a precise budget planning framework that facilitates the reliable verification of probabilistic AI models.
📝 Abstract
Third-party challenge-response identity verification (TP-CRIV) enables an independent verifier to assess whether a claimant possesses a model identical to a remotely deployed model without directly accessing the reference model. However, for probabilistic AI models, repeated executions of the same query may produce different outputs and therefore different verification observations. This raises the question of how such stochastic evidence should be accumulated and how much evidence is required for reliable verification. In this work, we characterize statistical separability in TP-CRIV of probabilistic AI models. Specifically, we relate challenge-wise behavior of matching and non-matching provers to verification-level separability. The characterization explicitly describes how the numbers of independent challenges and repeated responses affect detection performance and enables the verification budget required for a target AUC to be estimated. We instantiate the proposed characterization for LLMs using open-ended challenges. The experiments demonstrate matching-non-matching separation, close agreement between theoretical and empirical AUCs, and consistent estimates of the minimum verification budgets. These results provide a statistical basis for relating probabilistic model behavior to verification-level separability and the evidence required for third-party verification.
Problem

Research questions and friction points this paper is trying to address.

Third-party identity verification
Probabilistic AI models
Statistical separability
Challenge-response verification
Innovation

Methods, ideas, or system contributions that make the work stand out.

TP-CRIV
Probabilistic AI Models
Statistical Separability
Large Language Models
Verification Budget
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
T
Teruki Sano
Graduate School of Information Sciences, Tohoku University, Japan
Minoru Kuribayashi
Minoru Kuribayashi
Tohoku University
Multimedia SecurityMultimedia ForensicsInformation SecurityImage Processing
M
Masao Sakai
Center for Data-driven Science and Artificial Intelligence, Tohoku University, Japan
S
Shuji Isobe
Center for Data-driven Science and Artificial Intelligence, Tohoku University, Japan
E
Eisuke Koizumi
Center for Data-driven Science and Artificial Intelligence, Tohoku University, Japan
Z
Zhang Zhang
Center for Data-driven Science and Artificial Intelligence, Tohoku University, Japan
S
Satoru Matsumoto
Center for Data-driven Science and Artificial Intelligence, Tohoku University, Japan