Explainability from Training with Applications to TCR-Epitope Prediction

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the black-box nature of deep learning models in scientific applications by introducing a novel "training process interpretability" paradigm that dynamically tracks the learning trajectories and evidence organization mechanisms of TCR-epitope prediction models. Methodologically, we construct the TCR-XAI2 benchmark dataset integrating experimental and predicted structures, and perform model-agnostic interpretability analyses leveraging CNN and Transformer architectures alongside AlphaFold3. Our findings reveal distinct learning dynamics across architectures and uncover dual-chain evidence conflicts, while elucidating the mitigating role of MHC information. Furthermore, this approach effectively distinguishes between model preferences for empirical versus predicted data. Collectively, this work provides a new perspective for understanding model evolution during training in computational immunology.
📝 Abstract
Deep learning models have achieved strong performance in artificial intelligence for science, yet their black-box nature limits our understanding of how they learn scientific tasks. Existing methods for interpretability provide limited insight into how models organize evidence and evolve during learning. We introduce explainability from training (EFT), a model-agnostic paradigm that traces model interpretation during training to explain why models rely on specific features and how they organize these features as predictive evidence. We apply EFT to four state-of-the-art T cell receptor (TCR)-epitope prediction models, TCR-SRIM, TULIP, MixTCRpred, and NetTCR-2.2, spanning post-hoc and interpret-by-design approaches as well as transformers and CNNs. To investigate how structural information affects model explanations, we introduce a benchmark, TCR-XAI2, containing 388 unique experimentally resolved TCR-epitope structures, complemented by structures predicted using AlphaFold3, Boltz-2, TCRModel2, tFold-TCR, and OpenFold3. Using EFT with TCR-XAI2, we demonstrate that (1) CNN and transformer models exhibit distinct learning trajectories; (2) TCR $α$ and $β$ evidence can conflict during learning, limiting the benefits of jointly modeling both chains, while MHC information mitigates this; and (3) real versus predicted structural data for TCR-epitope prediction exhibits distinct TCR and peptide feature preferences as well as differing trajectories of model certainty.
Problem

Research questions and friction points this paper is trying to address.

Explainability
Deep Learning Interpretability
TCR-Epitope Prediction
Black-box Models
Structural Data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Explainability from Training
TCR-Epitope Prediction
Model-Agnostic Interpretability
Structural Benchmark
Learning Trajectories