🤖 AI Summary
This study addresses the governance challenges organizations face when deploying AI, which stem from an “AI assessability gap”—the lack of sufficient evidence to support high-confidence decisions. The work establishes evidence adequacy as a distinct dimension of AI governance and introduces the concept of “assessability”: the system’s capacity to continuously generate, maintain, and update adequate evidence for governance decisions. It distinguishes between operational and investment certification mechanisms and develops a formal theoretical framework grounded in a confidence function Conf(D|E), integrating structural and causal evidence analysis. Within this framework, six key attributes of assessable evidence are rigorously defined. By providing both theoretical foundations and practical pathways to bridge the assessability gap, this research constitutes a prerequisite for effectively managing AI-related risks and sustainably realizing its value.
📝 Abstract
Organizations deploying AI face two fundamental governance challenges: managing AI risk and sustaining AI value. Both depend on evidence whose sufficiency cannot be taken for granted. We call the shared underlying challenge the AI Evaluability Gap: the condition in which organizations lack sufficient evidence to support high-confidence governance decisions regarding either risk or value.
We argue that this gap reflects a category error in current practice. Existing governance approaches focus primarily on properties of systems, such as safety, fairness, reliability, compliance, and value, while paying comparatively little attention to the evidentiary foundations required to justify decisions about those properties. We further argue that AI governance encompasses both operational decisions regarding whether a system may operate and investment decisions regarding whether it merits continued organizational resources.
To address this problem, we introduce Evaluability, defined as the capability of a system to generate, maintain, and renew evidence sufficient to support high-confidence governance decisions over time. We formalize governance decisions as functions of calibrated confidence Conf(D|E) and identify six properties of evaluable evidence: observability, attributability, intervenability, verifiability, calibration, and temporal validity.
The framework distinguishes Operational Certification, which relies primarily on structural evidence to justify deployment decisions, from Investment Certification, which relies primarily on causal evidence to justify continued resource allocation. We argue that evidence sufficiency is a missing layer of AI governance and that closing the AI Evaluability Gap is a prerequisite for both managing risk and sustaining value in AI-enabled organizations.