๐ค AI Summary
Inferring organic molecular structures from spectral data constitutes an ill-posed inverse problem due to the sparsity of individual spectra and the vastness of chemical space. This work proposes a โhypothesize-and-refineโ learning paradigm: first generating chemically plausible structural hypotheses by leveraging multimodal spectra and large-scale molecular priors, then iteratively refining them under mass spectrometry constraints. To this end, we introduce QM9SPIN, the first DFT-based multimodal spectral dataset incorporating J-coupling, DEPT, and spin interactions, and pioneer the integration of high-resolution mass spectrometry constraints into conditional generative models to enforce global compositional consistency. Our SpectroMol framework, featuring the MS-Mol2Mol mass-constrained generator, achieves 93.8% Top-1 accuracy on simulated benchmarks and demonstrates effective transfer to real-world scenarios with minimal experimental fine-tuning, further enhanced through mass-spectrometry-guided refinement.
๐ Abstract
Determining molecular structures from spectroscopic data remains fundamentally challenging because the inverse problem is intrinsically underdetermined: individual spectra are sparse, low-dimensional, and encode only partial structural evidence relative to the vast space of possible molecules. We address this challenge by formulating automated structure elucidation as a scalable hypothesis-refinement paradigm that tightly integrates spectral evidence with large-scale molecular priors. To supply structure-resolving NMR signals for multimodal learning, we construct \textbf{QM9SPIN}, a DFT-derived dataset comprising diverse 1D and 2D spectra, including J-coupling, DEPT experiments, and explicit spin--spin interactions. On this foundation, we introduce \textbf{SpectroMol}, a spectrum-to-structure model that proposes chemically valid molecular hypotheses conditioned on multimodal spectral inputs. Complementarily, we develop \textbf{MS-Mol2Mol}, a high-resolution mass-constrained molecular generator that integrates molecular formula, exact mass, and degree of unsaturation within a conditional generative prior trained on 400 million molecules, ensuring global compositional consistency and chemically realistic refinement. The integrated system achieves 93.8\% top-1 accuracy on the simulated benchmark, adapts effectively from simulated to experimental spectra with limited experimental fine-tuning, and further improves experimental predictions through mass-guided refinement, establishing a scalable route toward automated, data-driven organic structure elucidation.