ProtocolMatch: Protocol-Dependent Model Selection for Scientific Dynamics Forecasting

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the evaluation bias arising from protocol-agnostic model selection in scientific dynamics prediction by proposing the ProtocolMatch framework. Using quantum spin systems as a benchmark, this work systematically compares architectures such as recurrent networks and causal attention mechanisms across diverse observational and computational protocols. The results reveal that performance rankings dynamically invert with varying training scales, and demonstrate that physics-informed penalty terms fail to consistently improve accuracy. By pioneering a protocol-dependent paradigm for model selection, this research establishes that single-metric accuracy is insufficient for guiding practical deployment. Instead, it advocates for the joint evaluation of physical validity and out-of-distribution reliability to ensure robust real-world applicability.
📝 Abstract
Scientific dynamics forecasting is often framed as an architecture choice, although deployment is also determined by observed history, rollout feedback, compute budget, physical objective, and test distribution. We formulate protocol-dependent model selection and introduce ProtocolMatch, a compute-matched, validation-selected, and failure-preserving evaluation framework. On driven quantum-spin dynamics, we compare recurrent, patched-attention, causal-attention, and low-rank linear predictors across three independently generated datasets. The causal-attention--recurrence ordering reverses as the training set grows within a fixed two-spin task, while a linear predictor has the lowest mean error in the six-spin local-observable comparison. Restricting observed history worsens every refreshed-history view but improves every closed-loop view in the four-spin study. A latest-state MLP has lower error than persistence on every dataset under state refresh across all five cells, yet its closed-loop rank varies by system and includes finite explosive errors. Physical penalties improve targeted consistency without reliably improving prediction error, and in-distribution intervals lose most coverage after a driving-frequency shift. Thus scientific model selection should return a predictor with its protocol and report accuracy, physical validity, and shifted-distribution reliability separately.
Problem

Research questions and friction points this paper is trying to address.

scientific dynamics forecasting
protocol-dependent model selection
model evaluation
distribution shift
architecture comparison
Innovation

Methods, ideas, or system contributions that make the work stand out.

Protocol-dependent model selection
Scientific dynamics forecasting
Causal attention
Closed-loop evaluation
Distribution shift
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.