DISSOLVR: An Interpretable and Fast Framework for Aqueous and Organic Solubility Prediction

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of physical interpretability in deep learning-based solubility prediction by proposing a transparent predictive framework. Methodologically, it employs physically grounded molecular descriptors as the basis for feature mapping and integrates large language model (LLM)-driven post hoc explanation techniques to generate chemical narratives, thereby effectively overcoming the traditional black-box paradigm. Results demonstrate that this framework achieves predictive accuracy approaching experimental error limits while exhibiting superior out-of-distribution generalization capabilities, earning strong endorsement from domain experts. By reconciling high predictive accuracy with robust transparency, this work establishes a novel paradigm for applying interpretable artificial intelligence in chemistry.
📝 Abstract
High-fidelity solubility prediction is fundamental to pharmaceutical development and environmental partitioning, where accurate modeling must couple molecular structure with thermodynamic behavior across diverse chemical environments. However, recent advancements have been dominated by deep learning architectures that often sacrifice physical interpretability for predictive power. We challenge this trend by showing that state-of-the-art performance does not require such non-transparent architectures. To address this, we introduce DISSOLVR, a transparent framework for molecular solubility prediction. In addition, we perform a comprehensive literature review and a benchmarking study against various methods. We show that DISSOLVR approaches the aleatoric limit of experimental uncertainty and achieves OOD generalization through structural invariance, derived by mapping molecules to physically-grounded descriptors. Then, we present an LLM-assisted post-hoc explanation pipeline that bridges the gap between symbolic model artifacts and chemically grounded narratives. Finally, a comparative benchmark of a survey involving 22 expert chemists reveals that expert evaluators provide deep insights.
Problem

Research questions and friction points this paper is trying to address.

solubility prediction
physical interpretability
deep learning
molecular structure
out-of-distribution generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Solubility Prediction
Interpretable Framework
Out-of-Distribution Generalization
Physically-grounded Descriptors
LLM-assisted Explanation
🔎 Similar Papers
No similar papers found.
V
Vansh Ramani
Indian Institute of Technology, Delhi, India
H
Har Ashish Arora
Indian Institute of Technology, Delhi, India
D
Dhairya Kuchhal
Indian Institute of Technology, Delhi, India
Sayan Ranu
Sayan Ranu
IIT Delhi
Machine learning for graphs
Tarak Karmakar
Tarak Karmakar
Associate professor at IIT Delhi (Computational Chem. Biol. & Mat.)
Computational ChemistryMolecular dynamicsEnhanced samplingMachine learning