The Specification as Quality Gate: Three Hypotheses on AI-Assisted Code Review

📅 2026-03-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses a critical limitation in current AI-based code review systems: in the absence of executable specifications, they often fall into structural loops and exhibit correlated errors due to shared training distributions between generation and review models, making it difficult to verify whether code aligns with true intent. The paper proposes a three-tiered architecture—“specification-first, deterministic verification, AI review of residual issues”—and provides the first systematic demonstration that executable specifications can transform code review from a complex domain into a complex yet solvable one. It further clarifies that AI should focus specifically on structural and architectural flaws beyond specification coverage. Through deterministic verification, cross-model review, and targeted defect injection experiments, the study confirms that both intra- and inter-family large models exhibit error correlation without specification guidance, whereas a specification-driven architecture effectively isolates the AI review boundary and significantly enhances reliability.

Technology Category

Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageCognitive Modeling & Cognitive Systems: Agent ArchitecturesHumans and AI: Other Foundations of Human Computation & AI

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
The dominant industry response to AI-generated code quality problems is to deploy AI reviewers. This paper argues that this response is structurally circular when executable specifications are absent: without an external reference, both the generating agent and the reviewing agent reason from the same artefact, share the same training distribution, and exhibit correlated failures. The review checks code against itself, not against intent. Three hypotheses are developed. First, that correlated errors in homogeneous LLM pipelines echo rather than cancel, a claim supported by convergent empirical evidence from multiple 2025-2026 studies and by three small contrived experiments reported here. The first two experiments are same-family (Claude reviewing Claude-generated code); the third extends to a cross-family panel of four models from three families. All use a planted bug corpus rather than a natural defect sample; they are directional evidence, not a controlled demonstration. Second, that executable specifications perform a domain transition in the Cynefin sense, converting enabling constraints into governing constraints and moving the problem from the complex domain to the complicated domain, a transition that AI makes economically viable at scale. Third, that the defect classes lying outside the reach of executable specifications form a well-defined residual, which is the legitimate and bounded target for AI review. The combined argument implies an architecture: specifications first, deterministic verification pipeline second, AI review only for the structural and architectural residual. This is not a claim that AI review is valueless. It is a claim about what it is actually for, and about what happens when it is deployed without the foundation that makes it non-circular.
Problem

Research questions and friction points this paper is trying to address.

AI-assisted code review
executable specifications
correlated failures
code quality
specification as quality gate
Innovation

Methods, ideas, or system contributions that make the work stand out.

executable specifications
AI-assisted code review
correlated failures
Cynefin framework
deterministic verification
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Christo Zietsman
Independent Researcher