Beyond Correctness: Evaluating Semantic Knowledge in Cross-Table Transfer

πŸ“… 2026-09-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses a critical flaw in semantic ablation experiments for tabular learning, where the absence of effective control groups frequently leads to misattributing models’ sensitivity to semantic perturbations as predictive gains, thereby producing misleading evaluations. To resolve this, the work proposes a novel perspective that disentangles content sensitivity from predictive utility, constructing an evaluation framework grounded in appropriate reference baselines. It establishes principled guidelines for task-specific control group selection and validates the approach through controlled ablation analysis and multi-study benchmark auditing. The findings reveal pervasive misattribution across existing literature, demonstrating that only a minority of prior studies provide evidence of genuine predictive gains under rigorous isolation. Ultimately, this research furnishes a more methodologically rigorous foundation for evaluating semantic knowledge in tabular learning.
πŸ“ Abstract
Semantic knowledge is increasingly used to bridge heterogeneous schemas in tabular learning, but how much does that knowledge actually improve prediction? Studies in tabular learning commonly answer this question through semantic ablations that modify or suppress the supplied semantic knowledge. We show that these ablations can lead to misleading conclusions about predictive benefit: poor performance under altered semantics may be taken as evidence that the intended knowledge is beneficial. Across real and controlled experiments, altering semantic content can produce large performance differences even when the model gains little predictive benefit from having that semantic knowledge in the first place. To separate these effects, we distinguish two quantities: content sensitivity and predictive utility. Content sensitivity measures the change in performance when semantic content is altered, whereas predictive utility measures the benefit of the intended semantic knowledge relative to a suitable reference without that knowledge. This distinction motivates an evaluation framework in which the control is chosen according to the question being asked: altered controls assess sensitivity to semantic content, whereas claims that semantic knowledge improves prediction require a suitable reference. Even then, predictive utility is not fixed; it varies across suitable references and decreases when the reference can more easily recover the tested knowledge from other inputs or labeled examples. In a bounded audit of 25 semantic-ablation comparisons across nine studies, only one of 18 explicit predictive-utility claims is paired with a control that clearly isolates the tested semantic contribution. Together, these findings motivate a simple evaluation principle: semantic-ablation controls should be chosen and interpreted according to the question they are intended to answer.
Problem

Research questions and friction points this paper is trying to address.

tabular learning
cross-table transfer
semantic knowledge
semantic ablation
predictive utility
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semantic Ablation
Predictive Utility
Content Sensitivity
Cross-Table Transfer
Evaluation Framework
πŸ”Ž Similar Papers
No similar papers found.