Score
Design and build systems that parse and normalize raw author/developer strings (personal names, email addresses, and affiliation text), extract institution and location fields, and standardize organization names. Reconcile and link variant identities across records and commits into canonical person and organization profiles (including multi-affiliated cases and country mapping), producing deduplicated author/developer identities with attention to precision and recall.
This work addresses the pervasive identity ambiguity in global code repositories—where individual developers use multiple accounts or identical identifiers are reused by different contributors—by proposing a high-precision identity disambiguation method. The approach constructs a global identity mapping spanning 5.87 billion commits, leveraging non-transitive clustering, a refined identity classification scheme (good/bad/local/bot/partial), project context restoration, and cross-commit provenance tracing. It successfully consolidates 107 million raw author strings into 62.7 million canonical identities while mitigating over-merging artifacts such as “giant clusters.” Evaluated on the ALFAA gold-standard benchmark, the method achieves a recall of 0.70 and precision of 0.88. Notably, 73.5% of commits are attributed to multi-string identities, and coverage of human developers increases to 98.17%.
This work addresses the pervasive ambiguity and duplication in author identity strings within ultra-large-scale code commit datasets, where conventional approaches often produce oversized clusters containing irrelevant developers due to over-merging. The authors propose a collaborative disambiguation framework that integrates graph-cutting with edge-level classification: leveraging betweenness centrality for graph partitioning and training a high-precision edge classifier (AUC = 0.99) using GitHub’s “no-reply” email labels. Complementary mechanisms—including fingerprint-based anti-aliasing, node filtering gating, and attribute blocklisting—are incorporated to enhance precision. Evaluated on the World of Code dataset comprising approximately 107 million author strings, the method substantially mitigates over-merging, reducing the largest cluster size from 170,431 to under 7,000 and improving gold-standard recall from 0.44 to 0.70. It further outperforms existing privacy-preserving global disambiguation techniques on a separate set of 21 million distinct GitHub identities, achieving an unprecedented balance between precision and recall at billion-scale.
This work addresses the challenge of accurately disambiguating developer identities in open-source software projects, where rampant use of aliases severely undermines the reliability of organizational and logical coupling metrics. To tackle this issue, the authors propose a scalable, high-precision developer identity deduplication pipeline that, for the first time, integrates large language model (LLM)-assisted annotation with human verification to construct a large-scale, high-quality dataset of duplicate identities. Building upon this dataset, they systematically evaluate the trade-offs among accuracy, inference time, and energy consumption across a range of classical machine learning models. The study contributes both the publicly released dataset and comprehensive benchmarking results, offering practical guidance for selecting cost-effective and accurate deduplication strategies tailored to diverse application scenarios.
Low accuracy in institutional normalization of author affiliation strings—characterized by nested multi-institutional structures and noise—hampers bibliometric analysis and cross-knowledge-base interoperability. Method: We propose AffRo, an end-to-end framework that jointly addresses affiliation parsing and coreference resolution, integrating rule-enhanced named entity recognition (NER), hierarchical organizational coreference resolution, and context-aware matching ranking. Contribution/Results: We introduce AffRoDB, the first expert-annotated benchmark dataset for affiliation normalization, filling a critical gap in systematic evaluation. On diverse, real-world affiliation strings from multiple sources, AffRo achieves a 12.6% absolute F1-score improvement over state-of-the-art methods, significantly enhancing scholarly metadata quality and enabling robust interoperation of organizational identifiers across knowledge bases.
This work proposes a unified natural language processing framework to address key challenges in academic integrity, including plagiarism, content fabrication, and authorship verification. The framework integrates four core stylometric tasks: classification of human- versus machine-generated text, distinction between single- and multi-author documents, detection of authorship changes within multi-author texts, and identification of contributing authors in collaborative writing. The study introduces and publicly releases the first academic text dataset generated using Gemini under two distinct instruction settings—standard and strict—and systematically evaluates how prompting strategies affect detection performance. Experimental results demonstrate that texts produced under strict instructions are significantly more adversarial, thereby increasing the difficulty of accurate identification. The code and dataset are made openly available, establishing a new benchmark for research on academic integrity.
Author name disambiguation in academic search is often hindered by cross-source inconsistencies and error propagation, while reliance on manual annotation incurs prohibitive costs. This work proposes CrossND, a novel framework that, for the first time, leverages cross-source inconsistency as a corrective signal to enable fully automated and highly robust disambiguation. CrossND integrates data cleaning, probabilistic soft logic reasoning, and test-time scaling into a chained refinement pipeline, eliminating the need for expert-labeled training data. Evaluated on real-world datasets, the method significantly outperforms 17 strong baselines, demonstrating the efficacy of cross-source reasoning in enhancing both accuracy and robustness in author name disambiguation.
This study addresses the vulnerability in software supply chains where compromised maintainer accounts enable malicious code submissions, a risk exacerbated by the absence of continuous behavioral authentication of author identity. To mitigate this, the authors propose a novel approach based on fine-grained, patch-level code style analysis, introducing cross-modal stylometry for commit verification in open-world settings. Their method employs a fine-tuned Transformer model to jointly embed code diffs and commit messages into a shared cross-modal embedding space, integrated with a streaming anomaly detection algorithm within CI/CD pipelines. Requiring no retraining, the system enables real-time detection of forged commits, achieving a ROC AUC of 0.93 on Linux kernel data. It successfully retrospectively identified known attacks such as the PHP backdoor and ForceMemo/GlassWorm incidents, while limiting forged commits to only 0.8%–1% of the review queue, substantially reducing manual auditing overhead.
This study systematically investigates the generalization gap of existing code authorship attribution methods when applied beyond competitive programming contexts to real-world classroom assignments. While state-of-the-art approaches achieve strong performance on benchmark datasets such as Google Code Jam—attaining 70.7% Top-1 accuracy among 1,000 authors—their effectiveness collapses dramatically in educational settings, dropping to 0.2% and 0.06% on university course assignments, levels nearly equivalent to random guessing. Using pre-trained Transformer models like CodeBERT, we conduct comprehensive benchmarking across multiple sources, including Google Code Jam, Kaggle, and a newly curated dataset of student homework submissions. Our findings reveal a pervasive and previously underappreciated cross-domain performance cliff, highlighting significant practical limitations of current techniques in authentic educational scenarios.
This work addresses data contamination caused by irreversible entity merging and ontology misclassification based on name fragments in knowledge graph construction. The authors propose a “review-before-linking” mechanism featuring an identity-ladder strategy—leveraging identifiers, names, and type scopes—to enable controlled deduplication, alongside anchor-evidence constraints that govern multi-class ontology label assignment. This approach corrects the evidential asymmetry arising when names are treated as instance labels rather than type assertions. Integrated into a system combining automated merging, evidence validation, and a human review queue, the method was evaluated on a knowledge graph comprising 537,157 entities and 2,198,567 relations. It reduced role assignment errors from 36 to zero, requiring only 775 manual decisions to resolve 48,403 merge proposals, thereby significantly mitigating risks of over-merging and misclassification.