π€ AI Summary
This study investigates the reliability and limitations of large language models (LLMs) in automatically generating entity-relationship (ER) diagrams from complex natural language requirements. Employing prompt strategies including zero-shot, chain-of-thought (CoT), and CoT augmented with a verifier, the authors systematically evaluate three leading LLMs on their ability to extract entities, relationships, and attributes from textual descriptions and produce conceptually consistent ER diagrams. The results indicate that while models perform adequately on low-complexity specifications, their outputs frequently suffer from logical inconsistencies, semantic ambiguities, and failures to correctly express constraints as requirement complexity increases. The findings highlight fundamental shortcomings of current LLMs in high-stakes database modeling tasks and provide empirical evidence for the role of prompt engineering in structured conceptual modeling.
π Abstract
This article analyzes the use of Large Language Models (LLMs) as support for the conceptual modeling of relational databases through the automatic generation of Entity-Relationship (ER) diagrams from natural language requirements. The approach combines different language models with prompt engineering techniques to evaluate their ability to identify entities, relationships, and attributes in a conceptually consistent manner. The experimental evaluation involved three LLMs, each subjected to three prompting techniques (Zero-Shot, Chain of Thought, and Chain of Thought + Verifier), applied to the same requirements scenario with progressively increasing complexity. The generated diagrams were qualitatively analyzed through direct comparison with the textual requirements, considering the structural and semantic adherence of the modeled elements. The results indicate that, although LLMs show reasonable performance in less complex scenarios, their reliability decreases as the complexity of the requirements increases, with a rise in inconsistencies, ambiguities, and failures in representing constraints. These findings reinforce that, in their current state, LLMs are not sufficiently mature for reliable use in complex scenarios, and the cost of validation may offset the apparent productivity gains.