Score
Designs and implements annotation adjudication systems and procedures that resolve labeling disputes by specifying adjudication protocols, criteria, and dispute-resolution workflows. Builds training for adjudicators, documents resolution outcomes, and measures inter-annotator reliability to validate and refine the protocol.
In high-risk industrial applications, severe label noise critically degrades model performance and trustworthiness. To address this, we propose a knowledge graph (KG)-guided multi-LLM agent committee framework: it dynamically constructs sample-context KGs to orchestrate debate-style reasoning and consensus voting among multiple LLM agents for automated identification and precise correction of noisy labels. Our key innovation is a coverage-based neural-symbolic error correction logic, achieving 100% recall on structural errors and high-fidelity repair. On the AlleNoise subset, our method attains an F1 score of 0.99—substantially outperforming single-LLM baselines (0.48) and KG-free committee variants (0.59). This work establishes a novel, interpretable, and verifiable paradigm for data purification in high-reliability machine learning systems.
This study addresses the lack of a unified data and modeling framework in financial alternative dispute resolution (ADR), which hinders accurate prediction of settlement outcomes. To overcome this limitation, the authors integrate data from multiple Japanese ADR institutions and propose a functional-label-based annotation scheme to characterize dispute structures. Building upon this representation, they develop a multi-task learning model that jointly performs dispute classification and settlement prediction. The work introduces, for the first time, a cross-institutional shared representation of dispute structure and demonstrates its partial generalizability across diverse ADR domains. Experimental results show that incorporating structural information significantly enhances settlement prediction performance, and when combined with large language models, the approach achieves or surpasses state-of-the-art results across multiple domains.
This work investigates the application of large language models (LLMs) to judicial assistance, focusing on two representative dispute domains: automobile insurance claims and domain name disputes. To address the challenge of transforming unstructured legal narratives into actionable insights, we propose a structured dispute analysis framework. First, it identifies key dispute elements to convert raw textual statements into structured summaries. Second, it employs multi-level prompt engineering to generate three-tiered, interpretable outputs: (i) identification of the prevailing party, (ii) assessment of claim acceptability, and (iii) evaluation of argument strength. To our knowledge, this is the first systematic effort to integrate LLMs into dispute resolution support by jointly leveraging information extraction, abstractive summarization, and hierarchical reasoning—thereby enhancing both granularity and interpretability of adjudicative assistance. Extensive experiments across multiple domain-specific datasets demonstrate statistically significant improvements over strong baselines across all three core tasks.
This study addresses the lack of systematic annotation and visualization methods for legal argumentation structures in Chinese judicial judgments, which hinders computational analysis of legal reasoning. Drawing on legal reasoning theory, the work proposes the first fine-grained and operational annotation framework tailored to Chinese judgments. It formally defines four types of propositions, five categories of argumentative relations, formal representation rules for nested structures, and corresponding visualization conventions. A standardized annotation protocol and inter-annotator consistency control mechanism are also developed. The resulting comprehensive and reproducible annotation guideline establishes a reliable methodological and data foundation for legal argument mining, computational modeling of legal reasoning, and AI-assisted legal analysis.
NLP data quality assessment has long relied on inter-annotator agreement, overlooking intra-annotator consistency—the temporal stability of individual annotators’ judgments. This neglect challenges the implicit “gold label as ground truth” assumption. Method: We conduct exploratory repeated annotation experiments across major NLP datasets and quantify intra-annotator agreement using Cohen’s and Fleiss’ Kappa, complemented by qualitative perceptual analysis. Contribution/Results: We demonstrate that mainstream NLP datasets routinely omit intra-annotator consistency reporting; moreover, individual annotators exhibit significant temporal variability in labeling identical texts. We identify and disentangle the dual influence of textual ambiguity and subjectivity on annotation stability. Our work establishes intra-annotator agreement as a foundational data quality metric, providing both a methodological framework and concrete guidelines for constructing more robust, reproducible NLP datasets.
This work addresses the challenge posed by the lack of structured semantic representations in legal case records, which hinders the performance of downstream legal AI tasks. To overcome this limitation, the authors propose LeDA, a web-based annotation platform that supports dynamic label creation without requiring a predefined ontology. LeDA enables annotators to iteratively discover and define legal concepts during the annotation process, while incorporating collaborative multi-user annotation and an arbitration mechanism to resolve disagreements. The system was successfully deployed on judgments from the Supreme Court of India, where three annotators constructed a “bag-of-concepts” semantic representation. This representation effectively facilitates precedent retrieval and judgment prediction, significantly enhancing the structured understanding and semantic processing of legal texts.
This study addresses the overreliance on inter-annotator agreement in current data annotation practices, which often overlooks annotation’s capacity to capture conceptual validity as a measurement act. Treating annotation as a measurement process, the work identifies five root causes of annotation issues—errors, ambiguity, impossibility, subjectivity, and annotator identity—and develops a measurement theory–based framework for diagnosing and improving annotation quality. Drawing on a synthesis of 132 literature sources and 10 semi-structured interviews, the research systematically defines target constructs, designs annotation instruments, implements labeling procedures, and evaluates both reliability and validity. The resulting framework equips annotation teams with evaluation methods that transcend mere agreement metrics, thereby substantially strengthening the foundational quality of AI training data.