🤖 AI Summary
This study addresses the limitation of existing LLM annotation pipelines that rely solely on label agreement for quality assessment, which obscures critical failure modes related to authorization compliance, information completeness, and evidence sensitivity. To overcome this, we construct an agentic LLM pipeline tailored for neuroimaging data (COBRE/FBIRN) and propose a decoupled auditing framework that treats authorization, completeness, and dynamic response as independent dimensions, complemented by a JSON parsing control mechanism to ensure structured outputs. This work reveals the evaluation blind spots inherent in traditional single-metric agreement approaches. By rendering latent errors explicit, the proposed multidimensional framework facilitates precise human-in-the-loop routing, thereby establishing a new paradigm for high-reliability medical AI annotation.
📝 Abstract
Many groups now delegate label ontology and metadata harmonization to agentic LLM pipelines. We built one and audited it. Our aggregate scores looked healthy, but the pipeline kept failing in ways they did not show, so we set out to find what they hid. Label agreement asks whether a proposed label matches a reference. It does not ask whether the agent was entitled to propose it, whether the output was complete enough to act on, or whether the label moved when the evidence moved. We measured those three separately on COBRE and FBIRN, two schizophrenia and control neuroimaging cohorts from different consortia, and they come apart, from label agreement and from each other. Showing the agent an upstream proposal barely moves label agreement, 0.857 to 0.870, while agreement on the chosen action doubles, 0.409 to 0.830. Output that parses as JSON still drops a required field on 10% of one model's cases and 33% of the other's. And an agent that replays its first answer scores perfectly on original cases and zero once we change the evidence that decides them. Downstream, a row-order error that none of these metrics reports erases most of the diagnostic signal. So we measure these properties apart, pair each with a control, and gate commitment on the result, which makes failures visible and easy to route to a person. None of this prevents failure. Our reference labels are rule-derived, so agreement with them means consistency, not correctness. Code is available at https://github.com/amir-sbg/Label-Agreement-Does-Not-Measure-Authorization.