🤖 AI Summary
This study addresses the absence of third-party held verification in the remote deployment of AI models by proposing the TP-CRIV framework. To our knowledge, this work pioneers a third-party statistical verification mechanism that requires neither white-box access nor protocol cooperation. Under the constraints of black-box interaction and network isolation, the framework leverages fresh challenges, probabilistically controlled witness generation, and independent threshold calibration to achieve model identity determination without external assistance. Experimental evaluations on ten ImageNet-pretrained CNNs demonstrate that the proposed method yields clear separation between same-model and cross-model scenarios, enabling effective identity verification within a limited number of challenges.
📝 Abstract
Artificial intelligence (AI) models are increasingly deployed through remote services, making model misappropriation a growing concern. Existing approaches, including watermarking, fingerprinting, and model similarity analysis, primarily rely on predefined evidence or direct behavioral comparison and do not explicitly evaluate whether the claimant currently possesses and can utilize model-dependent information relevant to the claimed model identity.
In this paper, we propose Third-Party Challenge-Response Identity Verification (TP-CRIV) for AI models. TP-CRIV targets a third-party verification setting in which the verifier has neither white-box nor API access to the claimant's model, can interact with the suspicious deployed service only through its ordinary black-box inference interface, and does not require protocol-specific cooperation from the service provider. Under these constraints, the framework enables the verifier to obtain empirical evidence as to whether the claimant locally possesses a model satisfying a predeclared identity relative to the deployed model. Verification is conducted under fresh, previously undisclosed requirements and network isolation, so that the demonstrated capability cannot rely on online external assistance after challenge disclosure. The resulting evidence is interpreted relative to independently specified and calibrated matching and non-matching operating situations and is statistical rather than cryptographic. We instantiate TP-CRIV for CNN image classifiers using probability-control-based witness generation. Experiments on ten ImageNet-pretrained TorchVision models demonstrate clear same/cross-model separation and finite-challenge verification using independently calibrated thresholds.