White Box Evidence Packages for Policy Audit Reports

📅 2026-07-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of evaluating whether policy audit reports generated by large language models are substantiated by valid evidence. The authors develop a controlled evaluation framework that fixes policy provisions, scoring criteria, and the underlying model while systematically varying the evidence interface to produce structured reports. These reports are then assessed by human annotators across multiple dimensions: correctness, relevance to policy provisions, diagnostic value, and evidence misuse. Innovatively framing internal model evidence as an evidence design problem, the work proposes a hybrid evidence package format and reveals a critical risk: when causal grounding is insufficient, models may erroneously reuse superficially plausible but irrelevant internal labels as evidence. Experiments on 600 reports derived from 60 AGORA cases demonstrate that the hybrid evidence interface yields optimal performance, whereas control conditions with scrambled evidence correlations show that models can generate seemingly coherent reports citing internally fabricated yet substantively irrelevant evidence—highlighting significant governance risks.
📝 Abstract
As AI governance moves from benchmark scores toward auditable oversight, a central question is how reviewers can tell whether an LLM-generated audit report is actually supported by evidence. This paper studies that question in passage-anchored policy audits, where a report must interpret a given policy passage and cite evidence for its claims. We introduce a controlled evaluation framework that holds the passage, rubric, and auditor model fixed while changing only the evidence interface supplied to the auditor. Across 60 AGORA policy cases, we generate 600 structured reports under ten evidence conditions, including passage-based evidence, internal model evidence, a hybrid package, and a shuffled control that preserves evidence format while breaking case relevance. Five human reviewers evaluate the primary interfaces for correctness, passage grounding, diagnostic usefulness, and evidence misuse. The results show that internal evidence changes how reports cite and reason about evidence, but more internal citations do not by themselves make a report more valid. A white-box diagnostic explains the failure mode: causal localization is narrow, while reports readily reuse broader readable labels and token directions. The hybrid interface is the most useful on average, while the shuffled control exposes a key governance risk: reports can sound substantively plausible while citing irrelevant internal evidence. This study reframes internal model access as an evidence design problem for audit workflows, rather than as a guarantee of transparency.
Problem

Research questions and friction points this paper is trying to address.

AI governance
audit reports
evidence support
policy interpretation
transparency
Innovation

Methods, ideas, or system contributions that make the work stand out.

white-box evidence
policy audit
evidence grounding
LLM transparency
audit workflow
🔎 Similar Papers
No similar papers found.