author name disambiguation

Methods for resolving ambiguous personal name strings into distinct author identities using bibliographic metadata, affiliation, and other signals. Employed to assemble large-scale bibliographic datasets, link authors to institutions and demographics, and analyze research concentration and bias.

authornamedisambiguation

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Practical Author Name Disambiguation under Metadata Constraints: A Contrastive Learning Approach for Astronomy Literature

Nov 13, 2025
VA
Vicente Amado Olivo
🏛️ Michigan State University | GSI Helmholtzzentrum für Schwerionenforschung | Foundation Future Industries

To address author name ambiguity caused by metadata sparsity in large digital libraries (e.g., NASA/ADS), this paper proposes a low-metadata-dependency contrastive learning framework for author disambiguation. It formulates disambiguation as a similarity learning task and employs a Siamese neural network to jointly embed author names, paper titles, and abstracts—explicitly avoiding reliance on fragile metadata such as affiliations or journal names. As a key contribution, the authors introduce the first large-scale, ORCID-aligned benchmark dataset for astronomy, enabling rigorous evaluation in a domain with scarce ground-truth annotations. Experiments demonstrate state-of-the-art performance: 94% accuracy on pairwise disambiguation and >95% F1-score on clustering—substantially outperforming existing methods. Robustness is validated on real astronomical literature. The code, pre-trained models, and evaluation dataset are publicly released to advance author disambiguation research under low-resource conditions.

Disambiguating author identities in astronomy literature with limited metadataGrouping publications by researcher identity using minimal available informationSolving name ambiguity issues in large digital libraries like NASA/ADS

Evaluating authorship disambiguation quality through anomaly analysis on researchers' career transition

Dec 25, 2024
HZ
Huaxia Zhou
🏛️ Northwestern University | Cold Spring Harbor Laboratory

Large-scale evaluation of author name disambiguation quality in biomedical literature remains challenging due to the absence of scalable, ground-truth–free assessment methods. Method: We propose an unsupervised quality proxy—“abnormal authorship position upon first independent PI appointment” (e.g., appearing as last author in the inaugural year)—and analyze temporal career trajectories of 5.8 million researchers from OpenAlex. Contribution/Results: Over 60% exhibit this anomaly, strongly correlating with disambiguation errors. Statistical analysis reveals that missing institutional affiliations significantly increase anomaly rates, whereas ORCID presence substantially reduces them. Notably, pre-2010 female authors show systematically elevated anomaly rates, suggesting gender-biased disambiguation may have compromised earlier demographic conclusions. This work establishes, for the first time, a robust association between early-career authorship patterns and disambiguation fidelity, introducing a scalable, annotation-free paradigm for evaluating large-scale author identification systems.

Authorship AttributionBiomedical LiteratureLarge-scale Databases

This study addresses the limited effectiveness of existing author disambiguation methods when applied to Chinese names, particularly in their romanized (pinyin) form. The authors propose the first rule-based disambiguation framework capable of uniformly handling both Chinese characters and pinyin, integrating co-authorship networks, citation graphs, institutional affiliations, and content similarity to achieve script-agnostic disambiguation. Evaluated on a manually annotated dataset of 80 author name pairs, the method achieves F1 scores of 0.88 and 0.89 for pinyin and character-based names, respectively—substantially outperforming baseline approaches. The primary gain stems from improved recall, demonstrating that the framework significantly enhances disambiguation accuracy and applicability for scholarly data involving non-Latin scripts.

author disambiguationChinese namesname ambiguity

LEAD: LLM-enhanced Engine for Author Disambiguation

Nov 10, 2025
GT
G. Tuccari
🏛️ Institute of Cognitive Sciences and Technologies (ISTC) | National Research Council of Italy (CNR) | Research Institute on Sustainable Economic Growth (IRCrES) | University of Catania

This paper addresses the author name disambiguation (AND) challenge across heterogeneous academic databases—CercaUniversità and Scopus. We propose LEAD, a lightweight hybrid framework that jointly leverages semantic features extracted by large language models (LLMs) and structural signals from collaboration/citation networks—including bibliographic coupling, co-occurrence analysis, and label propagation—to build an efficient learning pipeline. Evaluated on 606 real-world ambiguous cases, LEAD achieves 96.7% F1-score and 95.7% accuracy, significantly outperforming single-modality baselines in cross-source matching precision and scalability. Its key contribution lies in the first principled integration of LLM-driven semantic understanding with graph-structured evidence, enabling high performance while reducing computational overhead. LEAD establishes a novel paradigm for integrating multi-source scholarly data and enabling robust bibliometric assessment.

Developing hybrid methods combining semantic and structural evidence for disambiguationLinking academic career records with author profiles for accurate identificationResolving author name ambiguity across bibliographic databases and academic registries

This work proposes a novel dataset discovery framework that leverages citation contexts from scientific papers to better capture the semantic intent behind research queries, addressing the limitations of existing dataset search engines that rely primarily on metadata and keyword matching and consequently suffer from low recall. By treating citation context as the core signal—combined with large-scale context extraction, large language model–guided pattern recognition, and provenance-preserving entity resolution—the approach significantly reduces dependence on incomplete or inconsistent metadata. Evaluated on eight computer science queries, the method achieves an average normalized recall of 47.47% (peaking at 81.82%), substantially outperforming Google Dataset Search and DataCite Commons. The framework’s novelty and practical utility have been affirmed by domain experts across multiple disciplines.

citation contextdataset discoverymetadata

Latest Papers

What's happening recently
View more

Author name ambiguity often leads to erroneous merging or splitting of nodes in co-authorship networks, distorting network topology and compromising analytical reliability. This study systematically quantifies, for the first time, the systematic biases introduced by commonly used initial-based heuristic disambiguation methods across a range of network metrics. By constructing a high-precision benchmark network and employing counterfactual analysis combined with stochastic perturbations to simulate varying levels of disambiguation error, the authors demonstrate that initial-based disambiguation substantially underestimates network size, overestimates connectivity and collaboration intensity among authors, and obscures the true fragmentation of research communities. These findings reveal the potentially misleading impact of prevailing disambiguation practices on the analysis of scientific collaboration networks.

author name ambiguitycoauthorship networksnetwork properties

Author name disambiguation in academic search is often hindered by cross-source inconsistencies and error propagation, while reliance on manual annotation incurs prohibitive costs. This work proposes CrossND, a novel framework that, for the first time, leverages cross-source inconsistency as a corrective signal to enable fully automated and highly robust disambiguation. CrossND integrates data cleaning, probabilistic soft logic reasoning, and test-time scaling into a chained refinement pipeline, eliminating the need for expert-labeled training data. Evaluated on real-world datasets, the method significantly outperforms 17 strong baselines, demonstrating the efficacy of cross-source reasoning in enhancing both accuracy and robustness in author name disambiguation.

academic searchauthor name disambiguationcross-source reasoning

This work addresses the challenge of accurately linking global research funding agencies to their supported scholarly publications to enable systematic analysis of research funding flows. To this end, we propose a multi-stage hybrid disambiguation framework that integrates lexical normalization, similarity-based clustering, rule-based matching, named entity recognition, and manual validation to map ambiguous funding statements to unique organizational identifiers. Our approach successfully standardizes 1.9 million distinct funding strings, achieving high recall and precision as validated through document-level comparison and expert review. By publicly releasing annotations of match types alongside unresolved cases, our method significantly enhances transparency, reproducibility, and the capacity for large-scale analysis of the global research funding landscape.

bibliometric linkagefunding acknowledgmentorganization disambiguation

This study addresses the limitations of existing academic platforms in providing fine-grained author metadata and geographic visualization, which are either unavailable or prohibitively costly. The authors propose an automated analytical framework that leverages a single Google Scholar user ID and integrates data from five sources—Google Scholar, OpenAlex, CrossRef, Semantic Scholar, and OpenStreetMap—through a five-stage pipeline for paper parsing, author disambiguation, and geocoding. Key innovations include a Unicode-resilient metadata parser, a two-stage institutional similarity–based disambiguation mechanism, and a city-level location repair method using OpenAlex. The system increases city-level geographic coverage for authors from 0% to approximately 60%, reduces h-index attribution errors by up to ninefold, and produces interactive HTML maps alongside structured analytical reports.

author disambiguationbibliometric toolscitation analysis

This work addresses the pervasive ambiguity and duplication in author identity strings within ultra-large-scale code commit datasets, where conventional approaches often produce oversized clusters containing irrelevant developers due to over-merging. The authors propose a collaborative disambiguation framework that integrates graph-cutting with edge-level classification: leveraging betweenness centrality for graph partitioning and training a high-precision edge classifier (AUC = 0.99) using GitHub’s “no-reply” email labels. Complementary mechanisms—including fingerprint-based anti-aliasing, node filtering gating, and attribute blocklisting—are incorporated to enhance precision. Evaluated on the World of Code dataset comprising approximately 107 million author strings, the method substantially mitigates over-merging, reducing the largest cluster size from 170,431 to under 7,000 and improving gold-standard recall from 0.44 to 0.70. It further outperforms existing privacy-preserving global disambiguation techniques on a separate set of 21 million distinct GitHub identities, achieving an unprecedented balance between precision and recall at billion-scale.

author identity disambiguationcode commitsentity resolution

Hot Scholars

JH

Julian Hough

Swansea University
DialogueDialogue SystemsDisfluencyNatural Language Processing
NM

Nicholas Micallef

Senior Lecturer in Computer Science, Swansea University
Human-Computer InteractionUsable Privacy & SecurityMisinformationSocial Media Analysis
DS

Deshan Sumanathilaka

PhD Candidate at Swansea University
NLPMachine Translation and TransliterationMLWSD
AZ

Amir Zeldes

Associate Professor of Computational Linguistics, Georgetown University
corpus linguisticscomputational linguisticsNLPdiscourse