π€ AI Summary
This study addresses the prohibitive cost of expert annotation and the uncertain reliability of large language model (LLM) judges in large-scale entity alignment (EA) evaluation by conducting the first systematic benchmark of LLMs for structured prediction assessment. Through perturbation bias diagnostics, counterfactual label flipping, and double-blind human verification, we uncover the mechanism by which label exposure induces anchoring bias and identify a "frontier model paradox." Accordingly, we propose a label-free evaluation protocol that restores discriminative performance to near-ceiling levels. Furthermore, this work open-sources the first biomedical EA benchmark (MeSHβSNOMED CT) alongside a reproducible auditing framework, providing both theoretical grounding and practical tooling for the reliable deployment of LLM-as-a-judge paradigms.
π Abstract
Entity Alignment (EA) identifies equivalent entities across knowledge graphs and is critical for knowledge base integration and ontology merging. Evaluating EA systems at scale requires expensive expert annotation, making systematic assessment across diverse domains practically infeasible. LLM-as-judge evaluation offers a potentially scalable alternative, yet its reliability for structured prediction tasks like EA remains unstudied. We present the first systematic benchmarking study across three frontier models, three datasets, and four EA systems, using perturbation bias diagnostics, meta-evaluation across all dataset-judge-prompt combinations, and counterfactual label-flip tests. We identify anchor bias, a failure mode in which judges invert discrimination when the system's decision label is visible. Label exposure causally collapses judge discrimination (J-ROC-AUC 0.12-0.87), while a label-free protocol recovers near-ceiling capability on distinctive-name datasets (0.93-1.00) and significant recovery on biomedical pairs (0.93-0.95). Counterfactual experiments confirm causality (FSR 53-99%) and reveal a frontier model paradox: stronger judges exhibit greater label sensitivity, not less. A blinded two-annotator human evaluation (102 pairs, Cohen's kappa=0.902) confirms this mechanism directly. We release the first biomedical EA benchmark (MeSH-SNOMED CT, 15K pairs) and a reproducible auditing framework for LLM judge reliability in EA. Code and data are available at https://github.com/vaibhavalakshmiravideshik/llm-as-a-judge-entity-alignment.