CHARTER: Auditing Reference Substitution in Hierarchical Compact-Evidence Evaluation for Computational Pathology

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the implicit shift in reference targets caused by candidate filtering during compact evidence evaluation in digital pathology, which distorts conclusions regarding model fidelity and strategy comparisons. To resolve this, we introduce the first reference-aware evaluation framework and propose CHARTER, an evaluation charter that transforms implicit selection into auditable specifications by explicitly declaring objectives, quantifying predictive shifts, and auditing conclusion stability. This approach effectively distinguishes genuine predictive preservation from spurious gains induced by reference changes. Experiments based on multiple instance learning and random seed auditing demonstrate that the proposed framework identifies four deterministic reversals in key comparisons and successfully rectifies prior ACMIL comparison results, thereby significantly enhancing evaluation reliability and transparency.
📝 Abstract
In digital pathology, compact evidence is often used to explain or audit predictions made by whole-slide image multiple instance learning models. In hierarchical compact-evidence pipelines, candidate filtering introduces a strategy-specific candidate-conditioned prediction alongside the original full-bag prediction. If the evaluation reference changes while the intended target remains the original full-bag prediction, however, not only can the measured fidelity of the same compact evidence change, but comparisons between competing candidate strategies can also change. To make this dependence explicit, we introduce CHARTER, a reference-aware evaluation charter that asks researchers to DECLARE the intended target and reference, QUANTIFY candidate-induced prediction shift, and AUDIT the stability of comparative conclusions. Across the 15 comparisons in our main five-seed Random-K audit, 4 showed determinate reversals; in a matched native-ranking stress test, the ACMIL comparison changed from REVERSED to PRESERVED. CHARTER turns otherwise implicit candidate-filtering and reference choices into an auditable evaluation specification, helping distinguish genuine preservation of the intended prediction from apparent gains induced by changing the prediction being explained.
Problem

Research questions and friction points this paper is trying to address.

computational pathology
compact evidence evaluation
reference substitution
multiple instance learning
fidelity auditing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Computational Pathology
Compact Evidence Evaluation
Reference Substitution Auditing
Multiple Instance Learning
Evaluation Framework
🔎 Similar Papers
No similar papers found.