Score
Designs and evaluates abstract representations of information—conceptual, logical, and physical models such as entity‑relationship diagrams, schemas, ontologies, and mappings—that specify entities, attributes, relationships, constraints, and normalization rules. Builds and reviews database, warehouse, document, or schema artifacts and analyzes data structures, integrity, semantics, and mappings to support efficient storage, querying, integration, and downstream processing.
Traditional Entity-Relationship (ER) models suffer from semantic staticity due to oversimplified abstractions and lack alignment with relational database implementations. Method: This paper proposes a reconstruction of the ER model based on the unified “thimac” (thing/machine) paradigm, leveraging the Thinging Machine (TM) framework to model entities, attributes, and relationships as executable semantic units. It embeds five fundamental actions—create, process, release, transfer, and receive—to restore neglected process semantics. Contribution/Results: We present the first complete mapping from classical ER models to dynamic thimac structures, preserving ER’s simplicity while achieving semantic completeness and direct compatibility with the relational model. Empirical evaluation confirms support for multi-granularity modeling—including legacy ER—and demonstrates significant improvements in conceptual expressiveness and technical implementability, enabling seamless integration between conceptual design and relational implementation.
Contemporary relational database management systems (RDBMSs) suffer from insufficient logical data independence, reducing them to passive storage layers incapable of supporting modern architectural innovation. This paper argues that the Entity-Relationship (ER) model must serve as the native abstraction layer of RDBMSs to overcome this limitation, and it provides the first systematic theoretical justification and empirical validation of the ER model’s necessity and feasibility for ensuring logical independence. Based on this insight, we design and implement ErbiumDB—a prototype system integrating metadata-driven schema management, declarative relational semantic modeling, and runtime relationship evolution. Experimental evaluation demonstrates that ER-based abstraction significantly enhances decoupling between application and storage layers, enabling flexible, semantics-aware data management. ErbiumDB establishes a novel paradigm for intelligent database architectures and delivers a rigorously validated, extensible prototype foundation for future research and development.
This study investigates the reliability and limitations of large language models (LLMs) in automatically generating entity-relationship (ER) diagrams from complex natural language requirements. Employing prompt strategies including zero-shot, chain-of-thought (CoT), and CoT augmented with a verifier, the authors systematically evaluate three leading LLMs on their ability to extract entities, relationships, and attributes from textual descriptions and produce conceptually consistent ER diagrams. The results indicate that while models perform adequately on low-complexity specifications, their outputs frequently suffer from logical inconsistencies, semantic ambiguities, and failures to correctly express constraints as requirement complexity increases. The findings highlight fundamental shortcomings of current LLMs in high-stakes database modeling tasks and provide empirical evidence for the role of prompt engineering in structured conceptual modeling.
Legacy systems written in COBOL, PL/I, or Assembly—common in banking and telecommunications—are often undocumented and lack original developers, hindering comprehension and modernization. Method: This paper proposes a multi-language, cross-platform, customizable framework for constructing software knowledge graphs and interactively defining architectural boundaries. It integrates static code analysis, data schema parsing, and custom ontology modeling to enable expert-guided, incremental analysis of source code and data architecture, automatically identifying business- and data-driven logical boundaries and visualizing cross-boundary dependencies. Contribution/Results: The framework introduces the first knowledge-graph-driven approach for progressive modernization path planning and impact analysis. Evaluated on two real-world industrial systems, it significantly improves system understanding efficiency and enhances the accuracy of modernization strategy design.
Traditional database models, grounded in the set-valued functor paradigm, lack native support for algebraic operations—such as numerical comparison and arithmetic—and exhibit a fundamental semantic and computational gap with programming languages. To address this, we propose an algebraic database model that systematically embeds multiple Lawvere theories into a unified categorical semantics framework, thereby coherently formalizing schemas, instances, schema transformations, and queries. Leveraging a proarrow equipment—a double-categorical structure—we integrate all model components, enabling direct expression and execution of algebraic operations (e.g., addition, order comparison) within data constraints and queries. This approach bridges the foundational disconnect between database theory and programming language semantics, yielding a verifiable algebraic semantics for databases and establishing computational completeness.
This work addresses the limitation of traditional database logical design, which overlooks the capacity of large language models (LLMs) to comprehend schema semantics, thereby constraining Text-to-SQL accuracy. For the first time, LLM-friendliness is incorporated into logical schema design through three semantic-preserving and composable transformation strategies: abstraction (+A), workload-aware partitioning (+P), and descriptive renaming (+R). The proposed approach is compatible with both supervised and zero-shot settings, yielding consistent improvements across multiple Text-to-SQL models. Evaluated on the BIRD-Union and Spider-Union benchmarks, the method achieves up to a 4.2% absolute gain in execution accuracy, significantly enhancing the mapping from natural language queries to executable SQL statements.
This study addresses the limited semantic transparency and poor comprehensibility of existing conceptual models, which stem from their reliance on low-level syntactic constructs to represent domain abstractions, thereby hindering effective system design and stakeholder communication. To overcome this, the paper proposes a language-agnostic abstract symbol engineering approach that identifies, formalizes, visualizes, and validates recurring syntactic configuration patterns, replacing them with high-level, semantically transparent abstract symbols. The method is instantiated as the DeCleaR extension to Dynamic Condition Response (DCR) graphs. Empirical evaluation demonstrates that DeCleaR significantly enhances perceived model quality, pragmatic quality, and user preference compared to standard DCR graphs.
This study addresses the frequent operationalization failures that arise when large language models generate analytical workflows, stemming from a semantic gap between user intent and system-executable actions. Through cross-domain empirical analysis across finance, human resources, and public safety, the authors manually examined 236 analytical intents and their automatically generated workflows, systematically identifying and categorizing five distinct semantic-level failure patterns: comparative anchoring, procedural reasoning, quantitative reasoning, role confusion, and policy anchoring. The findings reveal fundamental limitations in the semantic expressiveness of current data systems and provide both theoretical grounding and practical guidance for improving the reliability of agent-generated analytical workflows.