The Checking Problem: What must be true before AI ships in a regulated firm

📅 2026-07-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the stagnation of enterprise AI initiatives in regulated financial institutions due to the absence of quantifiable evaluation criteria. Focusing on six document-intensive workflows, it systematically compares AI system performance across four model families and three tool configurations, distinguishing between demonstration and production environments. For the first time, it links deployment feasibility with human review rates. The authors propose a production-grade evaluation framework encompassing accuracy, reproducibility, traceability, and informative confidence, integrating multi-model comparison, confidence signals, source citation, and self-verification mechanisms. Experiments reveal that 56.1% of the 72 evaluated configurations meet production readiness thresholds. Incorporating source citation and confidence estimation reduces human review requirements to 49%, and adding self-verification further lowers this to 44%, albeit at the cost of reduced error tolerance.
📝 Abstract
Enterprise AI programmes stall at a rate that is widely quoted and poorly explained. This paper measures the mechanism. Six document-heavy workflows of the kind performed daily in regulated financial services were run across four model families and three tool configurations, three times each, producing 5,093 scored output elements across 72 configurations. Each configuration was assessed twice: against a demonstration bar, being a single correct run on a single case, and against a production bar requiring sustained accuracy, reproducibility across repeats, verifiable attribution, and a confidence signal that carries information. 57 of 72 configurations cleared the demonstration bar and 32 cleared the production bar, a survival rate of 56.1%. The paper then computes the review burden each configuration imposes, estimated out of sample rather than with hindsight. A tool that states no confidence requires review of 100% of its output, because it offers a reviewer no basis for triage. Requiring the tool to cite its sources and state a confidence reduces that to 49% while holding the residual error tolerance in 17 of 20 configurations. Adding a self-verification pass costs 2.3 times the latency of the plain configuration, reaches 44%, and is the only configuration that fails to hold the error tolerance. The practical implication is that the value of an AI workflow is set less by how often it is right than by how much of it a human must still check, and that the second property is measurable and rarely measured.
Problem

Research questions and friction points this paper is trying to address.

AI deployment
regulatory compliance
review burden
production readiness
confidence calibration
Innovation

Methods, ideas, or system contributions that make the work stand out.

production readiness
review burden
confidence calibration
verifiable attribution
AI deployment
💼 Related Jobs
No related jobs found.
P
Prerit Ahuja
Independent Researcher