Support-Set Target Leakage in Relational Foundation Models during In-Context Learning: Model Dependence and Evaluation Reliability

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the evaluation distortion in relational foundation models during in-context learning, where target-derived features within support sets induce multi-hop information leakage. Leveraging the RelBench benchmark, we construct controlled leakage scenarios and, for the first time, systematically quantify leakage effects across varying hop counts and configurations by integrating integrated gradients, mutual information, and leave-one-column-out cross-validation. Our findings demonstrate that zero-hop and full leakage introduce the most severe biases, capable of reversing model ranking conclusions. Furthermore, we reveal that existing detection methods fail to fully recover unbiased evaluations. This work establishes a novel paradigm for the reliability analysis of relational in-context learning.
📝 Abstract
Relational in-context learning (ICL) uses labeled support examples and their linked relational context to predict labels for new queries. This creates a failure mode when target-derived features are present in the support context but unavailable for the query. We study this setting as support-set target leakage. We construct 20 controlled target-derived features that vary in signal fidelity, representation, semantic transparency, coverage, and zero-, one-, and two-hop relational placement, and evaluate them across 13 RelBench tasks and five relational ICL configurations that vary the ICL head, message-passing depth, pretraining cohort, or relational encoder architecture. We evaluate matched 0-hop, 1-hop, and 2-hop leakage settings, together with a Full leakage condition containing all 20 leaker columns. Within the tested configurations, target-table (0-hop) and Full leakage produce the largest aggregate deviations from clean evaluation, while higher-hop effects are often weaker, consistent with differences in effective exposure associated with temporal reachability, sampling, and aggregation fidelity. Leakage effects are strongly task- and model-dependent and can reverse relative conclusions between model variants even when aggregate changes are small. For leaker detection, we compare an Integrated Gradients (IG)-based screening method with mutual information (MI) and leave-one-column-out (LOCO) on a common Baseline subset. Ranking quality is strongest in the high-impact 0-hop and Full leakage conditions, but detector-based removal does not consistently restore the clean evaluation. A four-task rel-salt case study further shows the same evaluation concern with native-schema leakage candidates from the original relational schema. These results identify the support/query information boundary as an important component of reliable relational ICL evaluation.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Support-Set Target Leakage
Relational In-Context Learning
Foundation Models
Leakage Detection
Evaluation Reliability
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Roshan Reddy Upendra
SAP, Palo Alto, CA, USA
A
Alexandre Dorais
SAP, Montreal, Canada
J
Joe Meyer
SAP, Palo Alto, CA, USA
A
Andrew Pouret
SAP, Montreal, Canada
A
Anastasios Lambrianos Stappas
SAP, Montreal, Canada
D
Dinesh Katupputhur Ramprasath
SAP, Palo Alto, CA, USA
Viswanath Ganapathy
Viswanath Ganapathy
AI Research, LG Advanced AI Labs
Artificial IntelligenceMachine Learning5G and Signal Processing
T
Tom Palczewski
SAP, Palo Alto, CA, USA
Minghua Li
Minghua Li
SAP, Seattle, WA, USA