🤖 AI Summary
This study addresses the challenge in computational biology where validation data frequently deviates from deployment scenarios, rendering model selection unreliable. To overcome this, we propose a forward proxy metric based on the Hessian norm—combining its trace and maximum eigenvalue—that leverages network geometric properties to evaluate model generalization across molecular and protein tasks. Our analysis reveals a fundamental disconnect between generalization signals and model selection signals. We demonstrate that low geometric scores correspond to collapsed predictors and, counterintuitively, that augmenting validation sets can degrade model selection performance under specific distribution shifts. Ultimately, this work establishes a novel paradigm for reliable model selection in complex, real-world scenarios.
📝 Abstract
Model selection in computational biology often relies on validation data drawn from the training regime, even when deployment lies outside it. When validation no longer preserves which model is best, a natural alternative is to rank candidates using properties of the trained network itself. We test this idea using a novel, forward-only proxy motivated by the norm of the Hessian, alongside common Hessian measures, across molecular property, protein fitness, and drug-response tasks. Contrary to our hypothesis, geometry does not become more useful as validation Spearman correlation deteriorates: augmenting validation helps some shifts but significantly harms others. More surprisingly, the proxy still correlates with generalisation gap on most tasks even when Hessian trace and top-eigenvalue relationships are weak or reversed, yet this signal does not reliably identify the deployment-best model. A curvature bound need not preserve cross-model rankings, and low geometric scores can even favour collapsed predictors. Thus, a generalisation signal need not be a model-selection signal.