Score
Designs and implements systems and algorithms that map, match, and normalize identifier strings across datasets and namespaces to persistent canonical identifiers; builds resolution services, mapping tables, and disambiguation logic to detect and reconcile mismatched or ambiguous identifier mentions. Analyzes identifier spaces and persistent identifier system semantics (namespaces, redirects, versions) to construct reliable processes for resolving, normalizing, and linking entity mentions to stable IDs.
Closed-class words (e.g., prepositions, conjunctions, articles) in source code identifiers—grammatically essential in natural language yet systematically understudied in programming language research—lack empirical characterization and theoretical grounding. Method: We construct CCID, the first manually annotated dataset of 1,275 closed-class identifiers, and integrate extended syntactic pattern modeling, grounded theory coding, and statistical analysis to uncover how such words encode control flow, data transformation, temporal logic, and behavioral roles via part-of-speech sequences. Contribution/Results: We propose a syntax-pattern–based framework for identifier semantic analysis and empirically demonstrate strong correlations between high-frequency closed-class patterns and program behavior. This work fills a critical gap in programming linguistics by providing the first large-scale empirical study of closed-class words in identifiers, with implications for identifier naming assistance, code comprehension, and programming pedagogy.
Cross-platform heterogeneity and incompleteness of metadata for life science software (e.g., bio.tools, Bioconductor) impede reliable software identity resolution. Method: We propose an automated solution for constructing a FAIR-compliant unified reference repository. This includes the first systematic evaluation of instruction-tuned LLMs (LLaMA, Phi-3) for software entity disambiguation, and introduces a novel multi-model ensemble inference framework with consensus-driven confidence modeling to enhance decision robustness. Contribution/Results: Our approach achieves >92% precision on real-world cross-source data and releases a human-annotated gold-standard benchmark. The study identifies fundamental limitations of current LLMs in fine-grained semantic discrimination and cross-registry terminology alignment. It provides a reproducible methodology to support sustainable analysis and observatory development for research software.
This study addresses identifier name similarity-induced naming confusion and its adverse effects on code comprehension, maintainability, and developer collaboration. To address the lack of a systematic classification framework in prior work, we propose the first taxonomy of identifier name similarity, spanning semantic, orthographic, and contextual dimensions—designed for both theoretical rigor and practical scalability. Through empirical analysis of naming patterns across large-scale open-source projects, we identify six high-frequency similarity categories (e.g., spelling variants, abbreviation conflicts, semantic near-synonyms) and empirically validate their prevalence and detrimental impact in real-world codebases. The taxonomy provides a reusable theoretical foundation and methodological support for identifier naming quality assessment, static analysis tool design, and collaborative naming convention development.
This work addresses the limitations of existing semantic caching approaches, which rely on embedding similarity for answer reuse but lack formal modeling of authorization, version control, and request equivalence. The authors propose a mathematically grounded, controllable reuse mechanism based on quotient sets: dialogue requests are parsed into structured representations, and fine-grained reuse chains are constructed via three classes of identity relations—reading, parsing, and reuse. Centered on reuse identity relations, the framework derives linguistic-side quotient objects induced by controlled answer partitions and establishes a reuse architecture that is finitely terminating, policy-admissible, and computationally complete. By integrating precise denotational semantics, aggregative design operators, and closure-stable units, the framework guarantees full computability and goal consistency under an untrusted proposal layer, ensuring that pipeline outputs remain invariant along identity chains in deployment logs and that every reuse instance requires verification via an identity or applicability certificate.
This work addresses data contamination caused by irreversible entity merging and ontology misclassification based on name fragments in knowledge graph construction. The authors propose a “review-before-linking” mechanism featuring an identity-ladder strategy—leveraging identifiers, names, and type scopes—to enable controlled deduplication, alongside anchor-evidence constraints that govern multi-class ontology label assignment. This approach corrects the evidential asymmetry arising when names are treated as instance labels rather than type assertions. Integrated into a system combining automated merging, evidence validation, and a human review queue, the method was evaluated on a knowledge graph comprising 537,157 entities and 2,198,567 relations. It reduced role assignment errors from 36 to zero, requiring only 775 manual decisions to resolve 48,403 merge proposals, thereby significantly mitigating risks of over-merging and misclassification.
Existing Semantic-ID (SID) tokenizers lack a unified diagnostic interface, causing mapping flaws—such as coverage gaps, full-code aliasing, and weak semantic prefixes—to remain undetected until downstream training. This work proposes the first systematic diagnostic framework tailored for SID mappings, which enables pre-training analysis by defining an adapter contract that integrates item mappings, metadata, and generation trajectories. The approach decouples addressability from behavioral semantic prefix evaluation and introduces mapping-level probes—including utilization rate, aliasing rate, neighborhood alignment, popularity distribution, and structural cost—as well as dynamic trajectory hooks. Experiments reveal that GRID-style mappings exhibit an aliasing rate as high as 0.977, whereas ReSID and GAOQ show no aliasing; deterministic category prefixes achieve the strongest co-occurrence alignment (0.447), confirming prefix alignment as a viable signal for candidate exposure.
This work addresses the poor generalizability of existing self-service entity resolution methods on unseen datasets and their difficulty in balancing precision and recall, which often leads to error propagation and cascading erroneous merges. To overcome these limitations, the authors propose a robust self-service entity resolution system that integrates three key innovations: automated selection among multiple algorithms, decoupled optimization of precision and recall—enforcing precision through rule-based veto mechanisms while enhancing recall via diverse candidate generation—and non-transitive cross-cluster merge validation to prevent error propagation. Extensive experiments on six real-world benchmark datasets, ranging from 864 to 5 million records, demonstrate that the proposed approach significantly reduces incorrect merges, enhances practical utility, and substantially lowers the trial-and-error burden for practitioners.
This study addresses the unclear role of distribution alignment in domain-aware entity matching under low-resource and varying supervision conditions. The authors systematically evaluate the BEACON framework across diverse data constraints and algorithmic configurations, offering the first in-depth analysis of the factors influencing distribution alignment in budget-constrained settings. Through controlled experiments, they uncover the synergistic mechanism between domain information and distribution alignment, demonstrating its critical impact on matching performance. The findings provide empirical evidence and practical design guidance for optimizing entity matching systems in low-resource, multi-domain scenarios.
Weak schema constraints in knowledge graphs often lead to predicate redundancy, resulting in semantic duplication, hindered reuse, and degraded data quality. This work is the first to frame predicate redundancy as a core data quality issue and proposes a closed-loop governance framework encompassing detection, resolution, and prevention. The approach integrates automated techniques—such as embedding-based clustering—with human-in-the-loop validation and embeds this synergy into a crowdsourced knowledge graph evolution pipeline. By extending the SciKGDash platform with interactive review capabilities and support for predicate merging or deletion, the system enables semi-automated curation. Evaluation on ORKG reveals that up to 30% of predicates are redundant, primarily due to user behavior and interface design flaws, thereby demonstrating the effectiveness of the proposed human–machine collaborative strategy.