🤖 AI Summary
Molecular images in chemical patents and related literature are difficult to retrieve via text-based search due to the absence of machine-readable structural representations.
Method: This paper proposes a novel vision-based fingerprinting paradigm that bypasses molecular graph reconstruction. It employs learned instance segmentation to precisely localize functional groups and carbon skeletons, followed by substructure encoding to generate disentangled visual fingerprint embeddings—eliminating reliance on conventional OCR and structure recognition pipelines.
Contribution/Results: The method exhibits superior robustness to stylistic variations—including distorted, hand-drawn, and low-resolution chemical diagrams. Evaluated on a multi-source chemical image dataset, it achieves state-of-the-art performance in visual retrieval, outperforming existing approaches (e.g., OCSR and conventional fingerprint methods), particularly on complex images. Retrieval accuracy improvements are statistically significant across challenging scenarios. This work establishes an efficient, scalable framework for chemical image understanding, with direct implications for drug discovery and materials science.
📝 Abstract
Automatic extraction of chemical structures from scientific literature plays a crucial role in accelerating research across fields ranging from drug discovery to materials science. Patent documents, in particular, contain molecular information in visual form, which is often inaccessible through traditional text-based searches. In this work, we introduce SubGrapher, a method for the visual fingerprinting of chemical structure images. Unlike conventional Optical Chemical Structure Recognition (OCSR) models that attempt to reconstruct full molecular graphs, SubGrapher focuses on extracting molecular fingerprints directly from chemical structure images. Using learning-based instance segmentation, SubGrapher identifies functional groups and carbon backbones, constructing a substructure-based fingerprint that enables chemical structure retrieval. Our approach is evaluated against state-of-the-art OCSR and fingerprinting methods, demonstrating superior retrieval performance and robustness across diverse molecular depictions. The dataset, models, and code will be made publicly available.