🤖 AI Summary
This study addresses the performance degradation and high retraining costs of pretrained mass spectrometry prediction models under shifts in chemical space and acquisition conditions by proposing the SPARC framework. This work introduces a pioneering retrieval-guided test-time specialization strategy that adaptively calibrates models without requiring source data or query spectra, leveraging retrieved reference spectra alongside a reliability-aware consistency mechanism. The framework synergistically integrates retrieval-augmented generation, test-time training, and reconstruction-based reliability assessment. Experimental evaluations demonstrate that SPARC significantly improves cross-domain prediction accuracy on benchmarks such as MassSpecGym, validating its effectiveness for real-world deployment scenarios.
📝 Abstract
Tandem mass spectrum prediction supports compound identification across metabolomics, natural-product discovery, and environmental analysis. However, pretrained predictors often degrade under shifts in chemical space and acquisition conditions, while retraining domain-specific models from scratch is costly. We introduce SPARC, a retrieval-guided test-time specialization framework that adapts a pretrained predictor using a spectral reference library without accessing test-query spectra. For each target query, SPARC retrieves chemically related reference spectra to recalibrate fragment intensities within the learned fragmentation space. During Transfer, SPARC combines reference-guided spectral adaptation with reliability-aware consistency, using reconstruction behavior on retrieved spectra to selectively preserve trustworthy predictions during continual specialization. Across MassSpecGym, NPLIB1 and application-specific GNPS libraries, SPARC improves spectral prediction under multiple transfer settings. These results establish retrieval-guided test-time specialization as a practical strategy for extending pretrained MS/MS predictors to specific chemical and acquisition domains, with continual test-time training providing further refinement during deployment.