Adjudicator: Correcting Noisy Labels with a KG-Informed Council of LLM Agents

๐Ÿ“… 2025-12-05
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
In high-risk industrial applications, severe label noise critically degrades model performance and trustworthiness. To address this, we propose a knowledge graph (KG)-guided multi-LLM agent committee framework: it dynamically constructs sample-context KGs to orchestrate debate-style reasoning and consensus voting among multiple LLM agents for automated identification and precise correction of noisy labels. Our key innovation is a coverage-based neural-symbolic error correction logic, achieving 100% recall on structural errors and high-fidelity repair. On the AlleNoise subset, our method attains an F1 score of 0.99โ€”substantially outperforming single-LLM baselines (0.48) and KG-free committee variants (0.59). This work establishes a novel, interpretable, and verifiable paradigm for data purification in high-reliability machine learning systems.

Technology Category

Machine Learning: Multi-class/Multi-label Learning & Extreme ClassificationNatural Language Processing: Fact-Checking / Misinformation Detection (NLP Focus)Data Mining & Knowledge Management: Representing, Reasoning, and Using Provenance, Trust

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: LLM based quality controls for crowd workGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
๐Ÿ“ Abstract
The performance of production machine learning systems is fundamentally limited by the quality of their training data. In high-stakes industrial applications, noisy labels can degrade performance and erode user trust. This paper presents Adjudicator, a system that addresses the critical data mining challenge of automatically identifying and correcting label noise and has been validated for production deployment. Adjudicator models this as a neuro-symbolic task, first constructing a dynamic Knowledge Graph (KG) to unify item context. This KG then informs a"Council of Agents,"a novel multi-agent Large Language Model architecture where specialized agents debate and vote on a label's validity. We validate our system on a 1,000-item balanced subset of the AlleNoise benchmark. Our KG-informed model achieves a 0.99 F1-score, significantly outperforming a single-LLM baseline (0.48 F1) and a non-KG council (0.59 F1). Our analysis reveals this is due to a Precision, achieved by a novel override logic that uses the KG to perfectly identify complex, structural errors (complete Recall) -- a class of errors that baselines fail to find. This result demonstrates a robust and explainable system for automated, high-precision data verification, serving as a vital proof-of-concept for generating golden datasets in strictly governed industrial environments.
Problem

Research questions and friction points this paper is trying to address.

Automatically identifying and correcting noisy training labels
Improving data quality for high-stakes industrial machine learning
Generating golden datasets in strictly governed environments
Innovation

Methods, ideas, or system contributions that make the work stand out.

Constructs dynamic Knowledge Graph for item context unification
Uses multi-agent LLM council debating label validity
Implements KG-informed override logic for error identification
D
Doohee You
Trusted Engagement & AI Solutions, Trust & Safety, Google
S
Sundeep Paul
Trusted Engagement & AI Solutions, Trust & Safety, Google