Label Agreement Does Not Measure Authorization

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing LLM annotation pipelines that rely solely on label agreement for quality assessment, which obscures critical failure modes related to authorization compliance, information completeness, and evidence sensitivity. To overcome this, we construct an agentic LLM pipeline tailored for neuroimaging data (COBRE/FBIRN) and propose a decoupled auditing framework that treats authorization, completeness, and dynamic response as independent dimensions, complemented by a JSON parsing control mechanism to ensure structured outputs. This work reveals the evaluation blind spots inherent in traditional single-metric agreement approaches. By rendering latent errors explicit, the proposed multidimensional framework facilitates precise human-in-the-loop routing, thereby establishing a new paradigm for high-reliability medical AI annotation.
📝 Abstract
Many groups now delegate label ontology and metadata harmonization to agentic LLM pipelines. We built one and audited it. Our aggregate scores looked healthy, but the pipeline kept failing in ways they did not show, so we set out to find what they hid. Label agreement asks whether a proposed label matches a reference. It does not ask whether the agent was entitled to propose it, whether the output was complete enough to act on, or whether the label moved when the evidence moved. We measured those three separately on COBRE and FBIRN, two schizophrenia and control neuroimaging cohorts from different consortia, and they come apart, from label agreement and from each other. Showing the agent an upstream proposal barely moves label agreement, 0.857 to 0.870, while agreement on the chosen action doubles, 0.409 to 0.830. Output that parses as JSON still drops a required field on 10% of one model's cases and 33% of the other's. And an agent that replays its first answer scores perfectly on original cases and zero once we change the evidence that decides them. Downstream, a row-order error that none of these metrics reports erases most of the diagnostic signal. So we measure these properties apart, pair each with a control, and gate commitment on the result, which makes failures visible and easy to route to a person. None of this prevents failure. Our reference labels are rule-derived, so agreement with them means consistency, not correctness. Code is available at https://github.com/amir-sbg/Label-Agreement-Does-Not-Measure-Authorization.
Problem

Research questions and friction points this paper is trying to address.

Label Agreement
Agentic LLM Pipelines
Metadata Harmonization
Evaluation Metrics
Neuroimaging
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic LLM Pipelines
Label Agreement
Evidence Sensitivity
Output Completeness
Gating Commitment
A
Amir Sabbaghziarani
Joint Georgia State University / Georgia Institute of Technology / Emory University Center for Translational Research in Data Science and Neuroimaging
B
Bradley Thomas Baker
Joint Georgia State University / Georgia Institute of Technology / Emory University Center for Translational Research in Data Science and Neuroimaging
T
Theodore J. LaGrow
Georgia Institute of Technology
Sergey Plis
Sergey Plis
TReNDS center: GSU, Emory, and GATech
Machine learning in brain imaging and beyond