data modeling

Designs and evaluates abstract representations of information—conceptual, logical, and physical models such as entity‑relationship diagrams, schemas, ontologies, and mappings—that specify entities, attributes, relationships, constraints, and normalization rules. Builds and reviews database, warehouse, document, or schema artifacts and analyzes data structures, integrity, semantics, and mappings to support efficient storage, querying, integration, and downstream processing.

datamodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-2.43
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$182K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Conceptual Entity-Relationship Model: Underneath the Simplicity and Staticity

Mar 08, 2025
SA
Sabah Al-Fedaghi
🏛️ Kuwait University

Traditional Entity-Relationship (ER) models suffer from semantic staticity due to oversimplified abstractions and lack alignment with relational database implementations. Method: This paper proposes a reconstruction of the ER model based on the unified “thimac” (thing/machine) paradigm, leveraging the Thinging Machine (TM) framework to model entities, attributes, and relationships as executable semantic units. It embeds five fundamental actions—create, process, release, transfer, and receive—to restore neglected process semantics. Contribution/Results: We present the first complete mapping from classical ER models to dynamic thimac structures, preserving ER’s simplicity while achieving semantic completeness and direct compatibility with the relational model. Empirical evaluation confirms support for multi-granularity modeling—including legacy ER—and demonstrates significant improvements in conceptual expressiveness and technical implementability, enabling seamless integration between conceptual design and relational implementation.

Addresses oversimplification in ER models for database designExplores static ER simplicity's impact on technical implementationProposes dynamic TM modeling to enhance ER compatibility

Beyond Relations: A Case for Elevating to the Entity-Relationship Abstraction

May 06, 2025
AD
Amol Deshpande
🏛️ University of Maryland

Contemporary relational database management systems (RDBMSs) suffer from insufficient logical data independence, reducing them to passive storage layers incapable of supporting modern architectural innovation. This paper argues that the Entity-Relationship (ER) model must serve as the native abstraction layer of RDBMSs to overcome this limitation, and it provides the first systematic theoretical justification and empirical validation of the ER model’s necessity and feasibility for ensuring logical independence. Based on this insight, we design and implement ErbiumDB—a prototype system integrating metadata-driven schema management, declarative relational semantic modeling, and runtime relationship evolution. Experimental evaluation demonstrates that ER-based abstraction significantly enhances decoupling between application and storage layers, enabling flexible, semantics-aware data management. ErbiumDB establishes a novel paradigm for intelligent database architectures and delivers a rigorously validated, extensible prototype foundation for future research and development.

Addressing insufficient logical data independence in RDBMSAdvocating shift from relational to entity-relationship modelExploring innovation via prototype system ErbiumDB design

This study investigates the reliability and limitations of large language models (LLMs) in automatically generating entity-relationship (ER) diagrams from complex natural language requirements. Employing prompt strategies including zero-shot, chain-of-thought (CoT), and CoT augmented with a verifier, the authors systematically evaluate three leading LLMs on their ability to extract entities, relationships, and attributes from textual descriptions and produce conceptually consistent ER diagrams. The results indicate that while models perform adequately on low-complexity specifications, their outputs frequently suffer from logical inconsistencies, semantic ambiguities, and failures to correctly express constraints as requirement complexity increases. The findings highlight fundamental shortcomings of current LLMs in high-stakes database modeling tasks and provide empirical evidence for the role of prompt engineering in structured conceptual modeling.

Conceptual Database ModelingEntity-Relationship DiagramsLarge Language Models

Legacy systems written in COBOL, PL/I, or Assembly—common in banking and telecommunications—are often undocumented and lack original developers, hindering comprehension and modernization. Method: This paper proposes a multi-language, cross-platform, customizable framework for constructing software knowledge graphs and interactively defining architectural boundaries. It integrates static code analysis, data schema parsing, and custom ontology modeling to enable expert-guided, incremental analysis of source code and data architecture, automatically identifying business- and data-driven logical boundaries and visualizing cross-boundary dependencies. Contribution/Results: The framework introduces the first knowledge-graph-driven approach for progressive modernization path planning and impact analysis. Evaluated on two real-world industrial systems, it significantly improves system understanding efficiency and enhances the accuracy of modernization strategy design.

Analyzing legacy systems for modernization using knowledge graphsIdentifying logical boundaries in large, undocumented software systemsUnderstanding dependencies to assess impact of incremental changes

Algebraic Databases

Feb 10, 2016
PS
Patrick Schultz

Traditional database models, grounded in the set-valued functor paradigm, lack native support for algebraic operations—such as numerical comparison and arithmetic—and exhibit a fundamental semantic and computational gap with programming languages. To address this, we propose an algebraic database model that systematically embeds multiple Lawvere theories into a unified categorical semantics framework, thereby coherently formalizing schemas, instances, schema transformations, and queries. Leveraging a proarrow equipment—a double-categorical structure—we integrate all model components, enabling direct expression and execution of algebraic operations (e.g., addition, order comparison) within data constraints and queries. This approach bridges the foundational disconnect between database theory and programming language semantics, yielding a verifiable algebraic semantics for databases and establishing computational completeness.

Functional GapTraditional DatabasesValue Functor

Latest Papers

What's happening recently
View more

This work addresses the limitation of traditional database logical design, which overlooks the capacity of large language models (LLMs) to comprehend schema semantics, thereby constraining Text-to-SQL accuracy. For the first time, LLM-friendliness is incorporated into logical schema design through three semantic-preserving and composable transformation strategies: abstraction (+A), workload-aware partitioning (+P), and descriptive renaming (+R). The proposed approach is compatible with both supervised and zero-shot settings, yielding consistent improvements across multiple Text-to-SQL models. Evaluated on the BIRD-Union and Spider-Union benchmarks, the method achieves up to a 4.2% absolute gain in execution accuracy, significantly enhancing the mapping from natural language queries to executable SQL statements.

LLM-friendly schemalogical database designschema transformation

This study addresses the limited semantic transparency and poor comprehensibility of existing conceptual models, which stem from their reliance on low-level syntactic constructs to represent domain abstractions, thereby hindering effective system design and stakeholder communication. To overcome this, the paper proposes a language-agnostic abstract symbol engineering approach that identifies, formalizes, visualizes, and validates recurring syntactic configuration patterns, replacing them with high-level, semantically transparent abstract symbols. The method is instantiated as the DeCleaR extension to Dynamic Condition Response (DCR) graphs. Empirical evaluation demonstrates that DeCleaR significantly enhances perceived model quality, pragmatic quality, and user preference compared to standard DCR graphs.

abstract notationconceptual modelinglow-level constructs

This study addresses the frequent operationalization failures that arise when large language models generate analytical workflows, stemming from a semantic gap between user intent and system-executable actions. Through cross-domain empirical analysis across finance, human resources, and public safety, the authors manually examined 236 analytical intents and their automatically generated workflows, systematically identifying and categorizing five distinct semantic-level failure patterns: comparative anchoring, procedural reasoning, quantitative reasoning, role confusion, and policy anchoring. The findings reveal fundamental limitations in the semantic expressiveness of current data systems and provide both theoretical grounding and practical guidance for improving the reliability of agent-generated analytical workflows.

agentic data systemsanalytical workflowslarge language models

Hot Scholars

JC

Jordi Cabot

Head of the Software Engineering RDI Unit at Luxembourg Institute of Science and Technology (LIST)
software engineeringmodelingopen sourcelow-code
VG

Vivek Gupta

Assistant Professor of Computer Science, Arizona State University
Artificial IntelligenceNatural Language ProcessingLarge Language ModelsInformation Retrieval
JS

Jin Song Dong

Professor of Computer Science, National University of Singapore
Formal MethodsTrusted AISafe AIModel Checking
JP

Javier Pastor-Galindo

Assistant Professor, University of Murcia
AISocial Network AnalysisDisinformationCyber Threat Intelligence
RL

Rodrigo Laigner

University of Copenhagen
Data-intensive ApplicationsCloud Data ManagementEvent-based Systems