A Chase-based Approach to Consistent Answers of Analytic Queries in Star Schemas

📅 2025-05-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Computing consistent answers to analytical queries over star-schema data warehouses with missing values and inconsistent data is computationally challenging. Method: We propose a polynomial-time algorithm based on the Chase procedure for consistent query answering (CQA), extending CQA—previously limited to simple queries—to analytical queries (involving projection, join, and selection) over star schemas. Under the independence condition for conjunctive selections on non-key attributes, we rigorously prove and construct an exact polynomial-time solution. Our approach employs dependency logic modeling, semantic constraint propagation, and an optimized Chase execution strategy to precisely compute consistent answers for standard SQL-style analytical queries. Contribution/Results: This work breaks theoretical and algorithmic barriers in CQA for star-schema data warehouses and analytical workloads. Unlike prior approaches—which either support only basic queries or rely on approximations—our method guarantees exact, efficient consistent query answers for realistic analytical queries under well-defined semantic constraints.

Technology Category

Data Mining & Knowledge Management: Intelligent Query ProcessingConstraint Satisfaction and Optimization: Satisfiability Modulo TheoriesKnowledge Representation and Reasoning: Computational Complexity of Reasoning

Application Category

Graph Algorithms and Modeling for the Web: Querying, indexing, and retrieval in Web-related graphsWeb Mining and Content Analysis: Community question answeringSemantics and Knowledge: Methods, algorithms and applications for the development of semantic models, knowledge graphs and other forms of structured data models with machine-interpretable semantics
📝 Abstract
We present an approach to computing consistent answers to analytic queries in data warehouses operating under a star schema and possibly containing missing values and inconsistent data. Our approach is based on earlier work concerning consistent query answering for standard, non-analytic queries in multi-table databases. In that work we presented polynomial algorithms for computing either the exact consistent answer to a standard, non analytic query or bounds of the exact answer, depending on whether the query involves a selection condition or not. We extend this approach to computing exact consistent answers of analytic queries over star schemas, provided that the selection condition in the query involves no keys and satisfies the property of independency (i.e., the condition can be expressed as a conjunction of conditions each involving a single attribute). The main contributions of this paper are: (a) a polynomial algorithm for computing the exact consistent answer to a usual projection-selection-join query over a star schema under the above restrictions on the selection condition, and (b) showing that, under the same restrictions the exact consistent answer to an analytic query over a star schema can be computed in time polynomial in the size of the data warehouse.
Problem

Research questions and friction points this paper is trying to address.

Computing consistent answers for analytic queries in star schemas
Handling missing values and inconsistent data in data warehouses
Extending polynomial algorithms to analytic queries with restrictions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Chase-based approach for consistent query answers
Polynomial algorithm for star schema queries
Handles missing and inconsistent data efficiently
🔎 Similar Papers
2024-04-15Annual Meeting of the Association for Computational LinguisticsCitations: 4