🤖 AI Summary
This study addresses the lack of scalability and unified frameworks for discovering multi-attribute logical and functional dependencies in tabular data by proposing LDTool and HLDTool. Transcending the limitations of traditional pairwise relationships, this work introduces a novel hypergraph-guided search space reduction strategy. By integrating optimization algorithms with multidimensional visualization techniques, the proposed tools enable the unified extraction and analysis of both dependency types. Experimental results demonstrate that these tools significantly reduce computational overhead on high-dimensional real-world datasets, successfully supporting dependency discovery involving hundreds of features. Consequently, this work provides an efficient and scalable solution for complex data analysis tasks.
📝 Abstract
Understanding the structural relationships among attributes in tabular data is fundamental to machine learning and pattern recognition. While functional dependency (FD) discovery has been extensively studied, scalable discovery of logical dependencies (LDs), particularly as the number of attributes and dependency order increase, remains underexplored. These dependencies capture non-deterministic, condition-specific relationships among pairwise or multiple attributes. Furthermore, existing approaches do not provide a unified framework for extracting multi-attribute LDs and FDs. To address these limitations, we propose LDTool and HLDTool for extracting and visualizing multi-attribute LDs and FDs from tabular data. LDTool extends dependency discovery beyond pairwise relationships, while HLDTool enables scalable extraction through hypergraph-guided search-space reduction. Experiments on three simulated and eleven real-world datasets demonstrate that the proposed framework extracts meaningful LDs and FDs while improving scalability. LDTool recovers the same FDs as existing FD discovery methods with lower runtime in high-dimensional feature spaces, whereas HLDTool enables dependency discovery in datasets with hundreds of features. The proposed framework provides interpretable visualizations of dependency structures and supports applications in exploratory data analysis and the quantitative evaluation of synthetic tabular data.