🤖 AI Summary
This study addresses delta learning in scientific machine learning, demonstrating that conventional residual magnitude inadequately characterizes target learnability, as complex baselines often yield rough learning targets that impede model convergence. To overcome this limitation, this work establishes baseline complementarity as a core principle and introduces scale-normalized graph Dirichlet roughness as a pre-training diagnostic metric. By integrating molecular graph neural networks with physics-based baseline modeling, the proposed framework optimizes target design for enhanced learnability. The findings reveal that small residuals are not inherently easier to learn; however, incorporating semi-empirical baselines effectively reduces target roughness. Consequently, this approach significantly improves both predictive accuracy and generalization capabilities across in-domain and out-of-distribution settings.
📝 Abstract
In scientific machine learning, $Δ$-learning trains models on residual errors relative to physical baselines, assuming that more accurate baselines with smaller residual scales inherently improve downstream performance. Here, we demonstrate that residual scale alone is an insufficient heuristic for learnability. Evaluating molecular graph neural networks on total energy targets, we show that complex local descriptor baselines can yield small residual targets that are disproportionately rough within architecture-informed proxy spaces and harder to learn relative to their scale. Conversely, semi-empirical baseline reduces both scale and normalized roughness, improving in-domain and out-of-domain prediction. We introduce scale-normalized graph Dirichlet roughness ($D_{\text{IQR}}$) as a pre-training diagnostic for residual learnability and establish baseline complementarity as a core target-design principle, elevating target space formulation alongside model architecture as a key axis for scientific machine learning.