Empty Intersection: Provenance Coverage Rose to 98% and Neither Verification Decision Moved

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue in data provenance verification where system outputs are frequently misclassified as observed values, thereby compromising downstream decision-making. To mitigate this, we propose structured defense mechanisms encompassing row-level hierarchical tagging, enforced single-entry write points, and query filtering, providing the first empirical quantification of their impact on production validation judgments. Experimental results demonstrate that these mechanisms increase classification coverage from 36.1% to 98.4% and successfully intercept 3,070 unauthorized writes. However, critical validation decisions remain unchanged, revealing a structural deficiency wherein defensive interventions do not intersect with the data informing those decisions. This finding substantiates the ineffectiveness of conventional provenance strategies in specific operational contexts.
📝 Abstract
Two structural defenses for provenance, a grade on every row, so that a verification routine cannot mistake the system's own output for an observation, and a single write ingress, so that the grade is enforced rather than merely conventional, were measured against the production deployment that motivated them, over a frozen snapshot of 194,620 rows and the two verification decisions the snapshot supports. Neither reaches either decision. Both were prescribed by a companion paper, which diagnosed that deployment: its verification routines decided outcomes using values the system itself had written. Neither prescription is new: both are established practice in fields that do not cite one another, and no prior work measuring whether either changes a verdict was found, so what is offered here is the measurement and not the prescriptions. Filtering the verification queries by grade turns both decisions from pass to undetermined; widening the grade vocabulary raises classified coverage from 36.1% to 98.4%; a single ingress requiring a grade refuses 3,070 writes. None of the three gives either decision admissible input. The prescriptions do not fail at what they specify. Each is stated over the population and makes no reference to any decision, so neither says which rows a decision will read, and the rows each intervention repairs and the 32 rows the decisions read do not intersect. The intervention that changed the most rows shows the reach most plainly: all 121,296 rows it moved from unnameable to named fall outside both query windows. This paper reports the conditions, measured rather than designed, under which the decisions would have admissible input at all, and notes that the two decisions are blocked for different reasons.
Problem

Research questions and friction points this paper is trying to address.

data provenance
verification decisions
structural defenses
provenance coverage
admissible input
Innovation

Methods, ideas, or system contributions that make the work stand out.

Data Provenance
Structural Defense
Verification Decision
Write Ingress
Provenance Coverage
🔎 Similar Papers
No similar papers found.