🤖 AI Summary
This study addresses the numerical ill-conditioning commonly encountered in dictionary learning for dynamical equation discovery in systems biology, where highly correlated candidate functions degrade model identification accuracy. The work systematically investigates the impact of multicollinearity in sparse regression on biological dynamical modeling, revealing that even a small subset of terms can induce severe ill-conditioning. Through comparative analysis of orthogonal polynomial bases and monomial bases under varying data distributions—supported by condition number assessments and numerical experiments on benchmark systems biology models—the study demonstrates that orthogonal bases substantially improve conditioning, numerical stability, and model recovery accuracy only when their associated weight functions align with the data sampling distribution.
📝 Abstract
Data-driven discovery of governing equations from time-series data provides a powerful framework for understanding complex biological systems. Library-based approaches that use sparse regression over candidate functions have shown considerable promise, but they face a critical challenge when candidate functions become strongly correlated: numerical ill-conditioning. Poor or restricted sampling, together with particular choices of candidate libraries, can produce strong multicollinearity and numerical instability. In such cases, measurement noise may lead to widely different recovered models, obscuring the true underlying dynamics and hindering accurate system identification. Although sparse regularization promotes parsimonious solutions and can partially mitigate conditioning issues, strong correlations may persist, regularization may bias the recovered models, and the regression problem may remain highly sensitive to small perturbations in the data. We present a systematic analysis of how ill-conditioning affects sparse identification of biological dynamics using benchmark models from systems biology. We show that combinations involving as few as two or three terms can already exhibit strong multicollinearity and extremely large condition numbers. We further show that orthogonal polynomial bases do not consistently resolve ill-conditioning and can perform worse than monomial libraries when the data distribution deviates from the weight function associated with the orthogonal basis. Finally, we demonstrate that when data are sampled from distributions aligned with the appropriate weight functions corresponding to the orthogonal basis, numerical conditioning improves, and orthogonal polynomial bases can yield improved model recovery accuracy across two baseline models.