🤖 AI Summary
This work addresses the scarcity of critical solubility parameter data that hinders the substitution of conventional solvents with greener alternatives. The authors propose a transfer learning approach based on a pretrained quantum chemistry Transformer foundation model, incorporating—for the first time—an uncertainty quantification mechanism to enable highly accurate prediction of Hansen solubility parameters and Gutmann donor–acceptor numbers from minimal labeled data. The resulting customizable and easily deployable screening tool substantially expands the coverage of available solvent data, successfully rediscovering known green solvents while also identifying novel candidate molecules. Experimental validation confirms the method’s practical utility in sustainable chemistry, and the team further releases an open-source high-throughput screening platform to facilitate community adoption and advancement.
📝 Abstract
Accurate prediction of solubility remains a central challenge across materials science and sustainable chemistry. In particular due to emerging technologies like organic and hybrid photovoltaics, batteries, and catalysis, solvent usage is expected to increase significantly within the coming years. Therefore, substituting solvents with greener alternatives is vital. This is where machine learning can have substantial impact. However, the limited data on critical parameters of solubility significantly constraints machine learning efficacy. In this work, we transfer a pre-trained foundational model on QM9 targets to our application with minimal data requirements. Additionally, the pipeline integrates uncertainty quantification, allowing the user to gauge the confidence of the predictions. As baseline, we succeed in predicting the Hansen solubility parameters and Dielectric Constant for which extensive databases exist. Importantly, we achieve high model performance on additional targets, such as Gutmann Donor and Acceptor numbers, where the available data is extremely limited. Overall, we augment data on solubility descriptors by orders of magnitude with high quality predictions. For effective dissemination, we deploy easy-to-use, easily integrateable with high throughput labs, customizable tool for ranking and screening possible solvent substitutes. Finally, we rediscovered known green solvent alternatives and proposed new candidates proving its relevance for finding eco-friendly solvents.