🤖 AI Summary
This study addresses the unfair comparisons and reproducibility challenges in existing link discovery methods caused by inconsistent benchmarks. To this end, we propose a unified evaluation framework spanning semantic, relational, and hybrid scenarios. By introducing composable structural and semantic perturbation strategies to generate reliable ground truth, alongside a source-specific validation mechanism, this work systematically evaluates diverse approaches, including set-based, feature-based, learning-based, and large language model embedding methods. Furthermore, the processed datasets, ground truth annotations, and perturbation generation pipelines are open-sourced. Collectively, this project establishes a fair, reproducible, and standardized evaluation foundation for the link discovery community.
📝 Abstract
Join discovery aims to identify tables from large data repositories that can augment a query table with complementary information, enabling downstream tasks such as data exploration, feature engineering, and business intelligence. Although numerous join discovery methods have been proposed, existing studies rely on method-specific benchmark construction, making reproducible and fair comparison difficult. We present TabJoinBench, a benchmark for evaluating join discovery methods across semantic, relational, and hybrid data lake scenarios. TabJoinBench constructs query-candidate pairs using source-specific validation strategies, systematically introduces structural, representation, and semantic changes through composable perturbations while preserving reliable ground truth. We evaluate representative join discovery methods spanning set-based, feature-based, and learned approaches, together with general-purpose language-model embedding baselines, and publicly release the processed datasets, ground-truth annotations, and generation pipeline to facilitate reproducible evaluation and future research.