🤖 AI Summary
This study addresses the high latency caused by over-reliance on generative models in schema matching, as well as the performance degradation induced by unconditional refinement. To tackle these issues, we propose JevNexus, a framework that integrates typed pairwise decisions with dual evidence from schemas and instances. Furthermore, it introduces a dynamic gating mechanism to enable conditional list refinement, which is triggered only upon decision discrepancies—requiring refinement for merely 5.6% of columns—thereby effectively mitigating the adverse effects of unconditional refinement. Experimental results demonstrate that JevNexus preserves high-accuracy metrics, including an MRR of 0.930 and Hits@1 of 0.909, while reducing average latency from 123 seconds to 16 seconds. This 7.7-fold efficiency improvement achieves an optimal balance between matching accuracy and computational overhead.
📝 Abstract
Schema matching increasingly uses generative language models to rerank retrieved column candidates, although the underlying task is a bounded correspondence decision. We present JevNexus, which combines typed pairwise decisions with schema/instance evidence and invokes listwise refinement only when the evidence disagrees and the fused margin is small. The evaluation covers 561 cases from six benchmark families. JevNexus obtains dataset-macro MRR and Hits@1 of 0.930 and 0.909, compared with 0.926 and 0.903 for Magneto, while reducing mean latency from 123.452 to 15.929 seconds (7.750). Paired analysis finds no statistically significant difference in either MRR or Hits@1. The gate invokes listwise refinement for only 5.665% of source columns and avoids the degradation caused by unconditional refinement. Code and experimental artifacts are available at https://github.com/RazeenLI/JevNexus.