QuanReview: Offline, Auditable Reconciliation of Human and LLM Span Annotations

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high cost and reliability challenges of employing large language models (LLMs) for structured annotation by proposing a dual-annotation-stream framework based on character-level alignment. The method automatically resolves unambiguous cases while routing conflicts to a browser-based interface for human adjudication. By integrating offline auditability, explicit logging strategies, and document- and span-level consistency computation, the framework supports field-level hybrid construction and direct export in original formats. Evaluated on a humanitarian benchmark, the system autonomously merges 8% of documents and precisely identifies 3,131 conflicts, substantially enhancing both review efficiency and result trustworthiness in human–machine collaborative annotation workflows.
📝 Abstract
Structured span annotations, such as quantities with their units, uncertainty modifiers, and event classes, are expensive to create and hard to keep trustworthy once language models enter the loop. We present QuanReview, an open-source system for auditing and correcting such annotation layers. QuanReview aligns two annotation streams over the same documents at character level, resolves unambiguous cases by an explicit and logged policy, and routes candidate conflicts to a browser-based adjudication interface where reviewers accept either side, build field-level hybrids, or flag items for re-annotation. A campaign manager assigns documents to multiple annotators with configurable redundancy, computes agreement at document and span level, auto-merges unanimous documents, and exports the corrected layer in the original file format, so that it can replace the original annotation files directly. Applied to a 4,457-record humanitarian benchmark and an LLM extraction stream, the system fully auto-merged 8% of documents, applied automatic policy decisions to a further 1,513 records, and concentrated human attention on 3,131 candidate conflicts, a mean of 5.4 per reviewed document.
Problem

Research questions and friction points this paper is trying to address.

span annotation
annotation reconciliation
LLM extraction
data auditing
structured annotation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Span Annotation Reconciliation
Auditable Alignment
Human-LLM Adjudication
Campaign Manager
Structured Annotations
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Matteo Musacchio
Universidad de San Andrés, Buenos Aires, Argentina
J
Juan Cruz Giner Pulero
Universidad de San Andrés, Buenos Aires, Argentina
I
Isabel Castañeda
Universidad de San Andrés, Buenos Aires, Argentina
N
Naomi Couriel
Universidad de San Andrés, Buenos Aires, Argentina
Yelena Mejova
Yelena Mejova
Senior Research Scientist, ISI Foundation
computational social sciencedigital epidemiologypublic healthsocial mediaculture
M
Mariano G. Beiró
Universidad de San Andrés, Buenos Aires, Argentina; CONICET, Buenos Aires, Argentina
Kyriaki Kalimeri
Kyriaki Kalimeri
Senior Researcher @ UNICEF, Researcher @ ISI Foundation
Computational Social ScienceMachine LearningHumanitarian AIData BiasesFairness