Tytan: Interactive Neurosymbolic Construction of Analytic Semantic Schemas from Relational Data

📅 2026-08-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the knowledge acquisition bottleneck inherent in traditional, manually crafted analytical semantic schemas, which are difficult to scale and heavily reliant on domain experts. The paper introduces the first end-to-end framework that integrates symbolic logic with large language models to automatically generate semantic schemas from relational databases, resolving ambiguities through natural language interactions to clarify user intent. By synergistically combining database schema parsing, symbolic reasoning, and semantic understanding, the method achieves 100% coverage of entities and attributes, 100% query executability, and semantic role accuracy ranging from 92% to 100% across seven benchmark domains. In blind evaluations, it successfully reconstructs the complete entity structure of a ten-table database without prior exposure.
📝 Abstract
From natural-language query interfaces to automated report generation, data analysis tools need a description of the data: the real-world entities it contains, which columns function as measures or identifiers, and how tables connect into units of analysis. Today, this semantic layer is usually written by hand. This is a knowledge-acquisition bottleneck that limits the scalability of analytic systems, keeps non-technical users dependent on experts, and is itself error-prone. We present TYTAN, a system for automatically constructing an analytic semantic schema from a relational database and, when available, a short user-provided description. TYTAN combines symbolic analysis of the database with LLM-based semantic inference for entity proposal, role assignment, and naming. When the evidence leaves a decision ambiguous, TYTAN asks the user a targeted natural-language question. We evaluate TYTAN on eight databases spanning real-world and benchmark domains along the three axes that define a schema's functional utility: (i) coverage, are all important entities and features captured?; (ii) retrieval correctness, do the schema's instructions actually reach the data; and (iii) characterization accuracy, are semantic types correct? Across the seven reference domains, TYTAN reaches every entity, attribute, and aggregable feature of the expert-corrected reference schemas (100% coverage). Additionally, 100% of its retrieval instructions execute correctly (1,678 of 1,678 self-generated claims), and semantic roles agree with the reference on 92-100% of matched attributes. Checking the underlying data showed the small disagreement is in the reference, not in TYTAN. On a held-out blind test (a live, ten-table database with no declared keys), TYTAN recovers the full entity structure with verified keys and satisfies 100% of the satisfiable expectations of five independent blind annotators.
Problem

Research questions and friction points this paper is trying to address.

semantic schema
relational data
knowledge acquisition bottleneck
data analysis
semantic layer
Innovation

Methods, ideas, or system contributions that make the work stand out.

neurosymbolic
semantic schema
relational data
LLM-based inference
interactive schema construction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Donna Hooshmand
Northwestern University
S
Shubham Shahi
Northwestern University
C
Cameron Barrie
Northwestern University
A
Abhratanu Dutta
Northwestern University
Marko Sterbentz
Marko Sterbentz
Northwestern University
Artificial IntelligenceNatural Language Processing
H
Harper Pack
Northwestern University
K
Kristian J. Hammond
Northwestern University