TabJoinBench: A Benchmark for Joinable Table Discovery

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unfair comparisons and reproducibility challenges in existing link discovery methods caused by inconsistent benchmarks. To this end, we propose a unified evaluation framework spanning semantic, relational, and hybrid scenarios. By introducing composable structural and semantic perturbation strategies to generate reliable ground truth, alongside a source-specific validation mechanism, this work systematically evaluates diverse approaches, including set-based, feature-based, learning-based, and large language model embedding methods. Furthermore, the processed datasets, ground truth annotations, and perturbation generation pipelines are open-sourced. Collectively, this project establishes a fair, reproducible, and standardized evaluation foundation for the link discovery community.
📝 Abstract
Join discovery aims to identify tables from large data repositories that can augment a query table with complementary information, enabling downstream tasks such as data exploration, feature engineering, and business intelligence. Although numerous join discovery methods have been proposed, existing studies rely on method-specific benchmark construction, making reproducible and fair comparison difficult. We present TabJoinBench, a benchmark for evaluating join discovery methods across semantic, relational, and hybrid data lake scenarios. TabJoinBench constructs query-candidate pairs using source-specific validation strategies, systematically introduces structural, representation, and semantic changes through composable perturbations while preserving reliable ground truth. We evaluate representative join discovery methods spanning set-based, feature-based, and learned approaches, together with general-purpose language-model embedding baselines, and publicly release the processed datasets, ground-truth annotations, and generation pipeline to facilitate reproducible evaluation and future research.
Problem

Research questions and friction points this paper is trying to address.

Join Discovery
Benchmark
Table Augmentation
Data Lake
Reproducibility
Innovation

Methods, ideas, or system contributions that make the work stand out.

Table Join Discovery
Benchmark
Data Lake
Composable Perturbations
Reproducible Evaluation
🔎 Similar Papers
2024-04-15Annual Meeting of the Association for Computational LinguisticsCitations: 4
2024-06-28IEEE International Conference on Data EngineeringCitations: 3