🤖 AI Summary
This work addresses the incompatibility between existing open-source cheminformatics tools and the scikit-learn ecosystem, which hinders unified and reusable molecular machine learning workflows. To bridge this gap, the authors introduce a Python library built on RDKit that fully adheres to the scikit-learn API specification. For the first time, core cheminformatics functionalities—including molecular fingerprinting, filters, similarity metrics, applicability domain estimation, and data splitting—are encapsulated within a consistent, composable interface. This design enables efficient computation and custom extensibility, supporting an end-to-end pipeline from SMILES inputs to deployable models. The proposed framework substantially enhances development efficiency, reproducibility, and system integration in molecular modeling, effectively reconciling cheminformatics with mainstream machine learning ecosystems.
📝 Abstract
We present scikit-fingerprints, a comprehensive, fully scikit-learn compatible library for molecular machine learning in Python, based on RDKit. Molecular fingerprints and related functionalities are workhorses of chemoinformatics, yet the widely used open-source frameworks are not compatible with the wider Python machine learning ecosystem based on scikit-learn conventions. scikit-fingerprints closes this gap, bringing molecular fingerprints, molecular filters, similarity and distance measures, applicability domain estimation, data splitting strategies, and more under a single, familiar interface. Scikit-learn compatibility means that an entire chemoinformatics workflow, from a raw SMILES string to a deployable model, can be assembled from composable building blocks and can reuse the mature tooling of the surrounding ecosystem. The underlying RDKit code makes it familiar and extensible for custom chemoinformatics use cases. We put a strong focus on unified interfaces, ease of use, computational efficiency, customization, and extensibility. scikit-fingerprints makes molecular machine learning faster to prototype, easier to reproduce, and simpler to deploy.