naming convention design

Designs and documents systematic rules, templates and patterns for assigning names to entities (identifiers, resources or artifacts), specifying syntax, structure, allowed characters, semantic encoding, and versioning. Creates naming strategies and governance artifacts — glossaries, templates, validators and change policies — and analyzes existing naming schemes for consistency, collisions, scalability and automation compatibility.

namingconventiondesign

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.36
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$187K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

10 Simple Rules for Improving Your Standardized Fields and Terms

Oct 21, 2025
RC
Rhiannon Cameron
🏛️ Simon Fraser University

Scientific data often suffers from poor discoverability, limited sharing, inefficient reuse, and high curation costs due to inadequate standardization of fields and terminology. To address these challenges, this paper proposes a FAIR-aligned standardization framework. Methodologically, it integrates structured vocabulary design, context-aware metadata modeling, and data homogenization strategies to systematically mitigate semantic noise and concept explosion. Crucially, it embeds ten actionable, principle-based rules into a dynamic, evolving data governance process. Empirical evaluation demonstrates that the framework significantly improves metadata quality and semantic consistency, reduces data management overhead, and enhances data findability, interoperability, and long-term reusability—thereby enabling robust, real-world implementation of the FAIR principles in scientific research settings.

Addressing challenges in standardizing research metadata vocabulariesOffering practical rules for FAIR-compliant metadata designProviding strategies to improve data findability and reusability

Although scientific data increasingly adhere to the FAIR principles and employ standardized identifiers, practical interoperability remains hindered by heterogeneity in identifier systems and data models. This work proposes and implements two synergistic tools—Babel and ORION—to bridge this gap. Babel constructs clusters of equivalent identifiers through mapping-based clustering and exposes them via a high-performance quantitative API, while ORION standardizes heterogeneous knowledge bases by aligning them to a community-governed common data model. Together, they systematically address the longstanding disconnect between the FAIR “Interoperable” principle and its real-world implementation. The integration of these tools has enabled the construction of a fully interoperable knowledge base, substantially enhancing cross-resource data integration and query capabilities. The resulting framework is publicly available.

Data ModelsFAIRIdentifier Schemas

From Instructions to ODRL Usage Policies: An Ontology Guided Approach

Jun 03, 2025
DM
Daham M. Mustafa
🏛️ Fraunhofer FIT | Universidad Privada Boliviana | RWTH Aachen University

This work addresses the need for automated digital rights policy generation in multi-institutional, culturally oriented trusted data spaces. We propose a large language model (LLM)-based natural language-to-ODRL policy mapping method. Our approach uniquely integrates the W3C ODRL ontology and its structured documentation as core components of prompt engineering to guide GPT-4 in generating high-fidelity, standards-compliant policies. Additionally, we introduce an ontology-adaptation heuristic tailored for knowledge graph construction to enhance semantic alignment. Evaluated on 12 culturally diverse use cases spanning varying complexity levels, our method achieves a policy generation accuracy of 91.95%, significantly outperforming existing baselines. The contribution lies in establishing a scalable, interpretable, and standards-aligned automation paradigm for open digital rights management—bridging natural language requirements with formal, machine-processable ODRL policies.

Automate ODRL policy generation from natural language instructionsEnhance policy accuracy using curated ontology documentationEvaluate approach in cultural dataspaces with 12 use cases

Closed-class words (e.g., prepositions, conjunctions, articles) in source code identifiers—grammatically essential in natural language yet systematically understudied in programming language research—lack empirical characterization and theoretical grounding. Method: We construct CCID, the first manually annotated dataset of 1,275 closed-class identifiers, and integrate extended syntactic pattern modeling, grounded theory coding, and statistical analysis to uncover how such words encode control flow, data transformation, temporal logic, and behavioral roles via part-of-speech sequences. Contribution/Results: We propose a syntax-pattern–based framework for identifier semantic analysis and empirically demonstrate strong correlations between high-frequency closed-class patterns and program behavior. This work fills a critical gap in programming linguistics by providing the first large-scale empirical study of closed-class words in identifiers, with implications for identifier naming assistance, code comprehension, and programming pedagogy.

Analyzes linguistic structure of identifier names with closed syntactic categoriesExplores relationship between closed-category grammar patterns and program behaviorInvestigates how developers encode behavior in source code via naming

Use of a Structured Knowledge Base Enhances Metadata Curation by Large Language Models

Apr 08, 2024
SS
Sowmya S. Sundaram
🏛️ Stanford University

This study addresses the low accuracy of large language models (LLMs) in FAIR-compliance validation of biosample metadata. We propose a structured-knowledge-guided prompting method, integrating the CEDAR template repository, domain-specific data dictionaries, and GPT-4 to construct a metadata standards-conformance verification framework—demonstrated on human lung cancer biosamples. Experimental results show that incorporating structured knowledge significantly improves field-level standards compliance from 79% to 97% (p < 0.01), providing the first empirical evidence that structured knowledge bases can overcome performance bottlenecks inherent to purely text-based LLM prompting in metadata governance. Our approach establishes a novel paradigm for automated, high-accuracy, and interpretable FAIR metadata quality control, enabling scalable, standards-aware curation of biomedical metadata.

Enhance metadata standards adherenceImprove metadata curation automationIntegrate structured knowledge with LLMs

Latest Papers

What's happening recently
View more

This work addresses data contamination caused by irreversible entity merging and ontology misclassification based on name fragments in knowledge graph construction. The authors propose a “review-before-linking” mechanism featuring an identity-ladder strategy—leveraging identifiers, names, and type scopes—to enable controlled deduplication, alongside anchor-evidence constraints that govern multi-class ontology label assignment. This approach corrects the evidential asymmetry arising when names are treated as instance labels rather than type assertions. Integrated into a system combining automated merging, evidence validation, and a human review queue, the method was evaluated on a knowledge graph comprising 537,157 entities and 2,198,567 relations. It reduced role assignment errors from 36 to zero, requiring only 775 manual decisions to resolve 48,403 merge proposals, thereby significantly mitigating risks of over-merging and misclassification.

data curationentity identityentity resolution

This study addresses the challenge in attributed graph schema design of whether repeatedly occurring descriptive attributes should be embedded within nodes or externalized as reusable metadata. Building upon Fifth Normal Form (5NF), the authors propose a principled decision framework that systematically identifies metadata candidates based on semantic criteria rather than mere repetition frequency. The approach classifies attributes into characteristic nodes, embedded properties, or borderline cases using five key principles: cross-element occurrence frequency, conceptual independence, lossless externalizability, reuse potential, and governance relevance. Empirical validation through a library domain case study and an entity classification task demonstrates that repetition alone is insufficient for externalization decisions—semantic judgment is essential. The proposed method significantly enhances the accuracy, consistency, and reusability of metadata modeling in graph-based systems.

embedded propertiesmetadataproperty graph schemas

This work addresses the challenges posed by the heterogeneous multimodal nature of enterprise policy documents, which often cause large language models to hallucinate, disrupt table structures, and lack end-to-end controllability—resulting in labor-intensive manual processing requiring 2–3 days per document. To overcome these limitations, the authors propose a governed multi-agent collaboration framework grounded in a shared, versioned rule repository. The framework integrates large language models (LLMs), vision-language models (VLMs), schema validation, and human-in-the-loop mechanisms through six specialized agents that collaboratively perform parsing, multimodal extraction, consistency verification, evaluation, iterative refinement, and personalized artifact generation, while ensuring full traceability across the pipeline. Evaluated on 120 real-world documents, the approach achieves a 96% success rate, automatically extracts 3,896 rules (71.4% auto-approved), produces 812 deployable artifacts, and reduces per-document processing time to 40–125 minutes.

enterprise guideline documentsgoverned workflowhallucinated content

Automatically generating YAML configuration files that are both structurally valid and compliant with multiple continuous integration (CI) service specifications remains a significant challenge, and the capabilities of current large language models (LLMs) on this task are not well understood. This work introduces DOC2CI, the first cross-CI benchmark dataset comprising 3,363 document–YAML pairs, and systematically evaluates 14 open-source models alongside GPT-series models. A novel failure taxonomy is proposed to uncover the root causes of model discrepancies, and this study provides the first empirical evidence that document similarity and structural validity constitute distinct optimization objectives. Experiments reveal that even the largest models achieve an Exact Match rate below 3.1%; while 97% of generated outputs are syntactically parseable, only 71% conform to the target service schema. Schema-guided post-hoc repair without additional training boosts structural validity to 94%, whereas fine-tuning improves document similarity at the expense of standalone structural correctness.

configuration generationContinuous IntegrationLLM

Hot Scholars

XQ

Xue Qin

Assistant Professor, Villanova University
Software EngineeringSoftware Testing
JS

John See

Professor, Heriot-Watt University Malaysia
Computer VisionImage ProcessingMultimediaArtificial Intelligence