🤖 AI Summary
This study addresses a critical challenge in retrieval-augmented generation (RAG) for tabular data: the misalignment between semantic relevance and actual answerability, where tables semantically similar to a query often lack sufficient evidence to answer it. The work systematically identifies this “semantic-answerability gap” and introduces TCR-Bench, a diagnostic benchmark built from homologous tables. To bridge this gap, the authors propose AAR, a lightweight, answerability-aware two-stage reranking method that leverages answerability judgment and row-column binding analysis—without requiring stronger underlying models—to substantially improve retrieval quality. Experimental results demonstrate that AAR boosts the top-1 target table retrieval rate from 18.2% to 57.4% and recovers question answering accuracy from 0.330 to 0.755, approaching oracle-level performance.
📝 Abstract
Tables are a critical knowledge source in retrieval-augmented generation (RAG), but a retrieved table may lack sufficient evidence to answer a query, a property we call answerability. While answerability broadly concerns whether a source or collection of sources contains sufficient evidence, retrieval models optimized for semantic relevance do not guarantee it even in the single-source case, creating a fundamental mismatch. To study this, we introduce TCR-Bench, a diagnostic benchmark for Table Content-level Answerability in RAG, built around sibling tables, i.e., tables with highly similar schemas but subtle content differences. On TCR-Bench, the dense retrievers we evaluate persistently exhibit a Semantic-Answerability Gap: they often retrieve the correct sibling group yet struggle to pinpoint the uniquely answerable table within it, dropping QA performance from 0.755 (oracle) to 0.330 (top-5 retrieved). Our analysis suggests this gap is associated with semantic accumulation, schema-level cue dependence, and weak row-column binding. As a diagnostic probe into the source of this gap, we test whether a lightweight two-stage pipeline, Answerability-Aware Reranking (AAR), applying direct query-table answerability judgment, can recover performance: it raises top-1 target retrieval from 18.2% to 57.4%, and this large gain is itself evidence that much of the observed failure reflects a missing answerability verification step, rather than an inherent limitation of model capacity alone.