Theoria: Rewrite-Acceptability Verification over Informal Reasoning States

📅 2026-07-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of providing auditable, high-assurance verification for AI-generated informal reasoning while maintaining broad coverage. It proposes a structured approach that reformulates solutions as typed state-transition sequences, where every step must be explicitly justified—via citations, computations, or given premises—and enforces a “change completeness” invariant to surface hidden assumptions and fabricated references. By integrating explicit justification licensing, human-readable proof traces, and an adversarial validation framework, the method enables transparent, traceable, and contestable reasoning. Empirical evaluation demonstrates strict certification accuracies of 91.4% on HLE-Verified Gold and 97.1% on GPQA Diamond, significantly outperforming existing monolithic LLM-based evaluators in detecting concealed premises and hallucinated citations.
📝 Abstract
When should an AI system's answer be trusted? Formal proof assistants offer certainty but cannot reach most of the problem distribution; scalar LLM judges offer coverage but produce opaque scores that cannot be audited after the fact and are subject to the same coherence issues as any LLM. We present Theoria, a verification architecture that closes this gap. A candidate solution is rewritten into a sequence of typed state transitions, each licensed by an explicit justification, whether that be a citation, computation, or problem-given fact, and every transition is independently auditable. The foundational invariant is completeness of change: every difference between consecutive proof states must be accounted for, so hidden premises surface as unlicensed mutations rather than passing silently. On HLE-Verified Gold (185 text-only expert problems), Theoria certifies 105 at 91.4% strict precision (Wilson 95% CI [84.5%, 95.4%]). Every certification produces a human readable proof trace in which each step can be independently challenged. Holistic LLM judges achieve comparable precision at matched coverage but fail on different problems (Jaccard 0.14-0.36), making the approaches complementary. On 95 adversarial poisoned proofs across 15 domains, structured judges catch 94.7% versus 83.2% for holistic judging (p= 0.0017). The overall 11.5 pp gap concentrates in hidden premises (90.6% vs. 62.5%, a 28 pp difference) and fabricated citations (100% vs. 90%), the error classes where the formal analysis predicts an advantage; performance is identical on arithmetic and theorem-misapplication errors, where no advantage is predicted. On GPQA Diamond (n= 65), certified precision is 97.1% (Wilson CI [85.1%, 99.5%]).
Problem

Research questions and friction points this paper is trying to address.

trustworthiness
verification
informal reasoning
auditability
AI reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

rewrite-acceptability verification
typed state transitions
completeness of change
auditable reasoning
hidden premise detection
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Michael Saldivar
Independent Researchers
B
Ben Slivinski
Independent Researchers