Data-Error Scaling in Machine Learning on Natural Discrete Combinatorial Mutation-prone Sets: Case Studies on Peptides and Small Molecules

📅 2024-05-08
🏛️ arXiv.org
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the scaling relationship between dataset size and prediction error for machine learning models operating in highly mutable discrete combinatorial spaces—such as proteins and small molecules—where conventional continuity assumptions break down. Method: We introduce a mutation-oriented data reordering strategy and a normalized learning curve analysis framework, integrating kernel ridge regression, synthetic multi-body-theoretic data, calibration plot clustering, and resampling techniques. Contribution/Results: We discover, for the first time, a “saturation–asymptotic” two-stage learning paradigm driven by mutational complexity, accompanied by a discontinuous drop in test error at a critical dataset size—a phase-transition-like phenomenon. Systematic validation on peptide–protein binding affinity and small-molecule solvation energy prediction tasks demonstrates that mutational complexity is the dominant factor governing learning efficiency and generalization performance, substantially outperforming conventional metrics such as sequence length or chemical diversity.

Technology Category

Machine Learning: Learning TheorySearch and Optimization: Mixed Discrete/Continuous SearchData Mining & Knowledge Management: Scalability, Parallel & Distributed Systems

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingWeb Mining and Content Analysis: Machine learning and data science for the WebSecurity and Privacy: Large-scale security measurements
📝 Abstract
We investigate trends in the data-error scaling behavior of machine learning (ML) models trained on discrete combinatorial spaces that are prone-to-mutation, such as proteins or organic small molecules. We trained and evaluated kernel ridge regression machines using variable amounts of computationally generated training data. Our synthetic datasets comprise i) two na""ive functions based on many-body theory; ii) binding energy estimates between a protein and a mutagenised peptide; and iii) solvation energies of two 6-heavy atom structural graphs. In contrast to typical data-error scaling, our results showed discontinuous monotonic phase transitions during learning, observed as rapid drops in the test error at particular thresholds of training data. We observed two learning regimes, which we call saturated and asymptotic decay, and found that they are conditioned by the level of complexity (i.e. number of mutations) enclosed in the training set. We show that during training on this class of problems, the predictions were clustered by the ML models employed in the calibration plots. Furthermore, we present an alternative strategy to normalize learning curves (LCs) and the concept of mutant based shuffling. This work has implications for machine learning on mutagenisable discrete spaces such as chemical properties or protein phenotype prediction, and improves basic understanding of concepts in statistical learning theory.
Problem

Research questions and friction points this paper is trying to address.

Investigating data-error scaling laws in mutation-prone combinatorial spaces
Analyzing discontinuous error phase transitions in ML training regimes
Developing normalization strategies for mutagenizable discrete space learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Kernel ridge regression on mutation-prone datasets
Mutant-based shuffling for learning curve normalization
Discontinuous phase transitions in data-error scaling laws
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
University of Basel | ETH Zurich | Swiss Nanoscience Institute | University of Toronto | Vector Institute | TU Berlin
V
Vanni Doffini
University of Basel, ETH Zurich, Swiss Nanoscience Institute
O
O. Anatole Von Lilienfeld
University of Toronto, Vector Institute, TU Berlin
M
Michael A. Nash
University of Basel, ETH Zurich, Swiss Nanoscience Institute