DIADA: Automatic Data Composition in Data Lakes

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of scattered attributes, high manual integration costs, and the absence of a universal organizational paradigm in data lakes by formally defining, for the first time, the task-agnostic data combination problem and proposing the DIADA system. Grounded in multivariate dependencies, DIADA employs predicate space mapping and inclusion lattice structure mining to design efficient and scalable combination algorithms. It further leverages independence assumption violation detection to evaluate association significance, thereby automatically reorganizing heterogeneous data lakes into semantically coherent datasets. Experimental results demonstrate that DIADA significantly reduces noise while enhancing statistical correlations, effectively improving both the reliability and efficiency of pattern discovery in large-scale environments.
📝 Abstract
Data lakes contain a plethora of attributes scattered across many tables that, when combined, provide enhanced assets for data analysis. Nonetheless, deciding which attributes belong together in meaningful relations remains a manual, per-task effort. Merging by joinability alone provides no guarantees regarding attribute relevance, while selecting features against a single target discards attributes useful to other tasks. To address this gap, we introduce the data composition problem: organizing a fragmented, heterogeneous lake into meaningful relations, agnostic of any particular analytical task so that the resulting organization can serve as a common foundation for diverse downstream analyses. We propose DIADA, a composition system that employs multivariate dependence as the criterion for assessing the meaningfulness of a relation and approximates it by hypothesizing independence among attributes and identifying those sets that violate this hypothesis. To do so, we map the attributes to a predicate space, forming a lattice under inclusion and mining those predicate sets that exhibit dependence among their constituents. We contribute a dedicated and scalable algorithm to effectively explore this space, outscaling classical algorithms for mining relationships, thus discovering dependencies that would otherwise be impractical to identify. We demonstrate that applying a single data composition process benefits diverse potential downstream tasks. This is the result of providing a subset of low-noise, statistically relevant attributes that increases the confidence that detected patterns are grounded in real relationships, thus preventing common modeling issues in large-scale environments.
Problem

Research questions and friction points this paper is trying to address.

Data Lakes
Data Composition
Attribute Relevance
Multivariate Dependence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Data Composition
Multivariate Dependence
Predicate Lattice Mining
Scalable Algorithm
Data Lakes
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Marc Maynou
Universitat Politècnica de Catalunya, Barcelona, Spain
A
Albert Martin
Universitat Politècnica de Catalunya, Barcelona, Spain
S
Sergi Nadal
Universitat Politècnica de Catalunya, Barcelona, Spain
Anna Queralt
Anna Queralt
Universitat Politècnica de Catalunya
Data managementData governanceCloud continuumAutomated reasoning
Oscar Romero
Oscar Romero
Universitat Politècnica de Catalunya, BarcelonaTech
Data ManagementData GovernanceData IntegrationBig DataData Science