🤖 AI Summary
This work addresses a critical limitation in existing cross-modal molecular representation methods, which often neglect the intrinsic structure of chemical space when aligning molecular structures with cellular phenotypes, leading to representation distortion and loss of structural information. To overcome this, the authors propose PhenMol, a novel framework that introduces structure-preserving constraints into multimodal molecular representation learning for the first time. PhenMol decouples shared and modality-specific representations, enabling integration of cellular phenotype data while preserving chemical neighborhood relationships through a dedicated molecular branch. The approach combines representation disentanglement, cross-modal alignment, and a structure-preserving loss, with ECFP4 fingerprints used for validation. Extensive experiments demonstrate that PhenMol achieves significant performance gains across 270 tasks, including bioactivity prediction, molecule–phenotype retrieval, and clinical trial outcome forecasting, while effectively mitigating structural distortion in the embedding space.
📝 Abstract
Phenotypic drug discovery enables the discovery of functional relationships between molecular structures and cellular responses. However, existing multimodal representation learning methods often optimize cross-modal alignment without considering the intrinsic organization of chemical space, resulting in distorted molecular representations and loss of structural information. We propose \textbf{PhenMol}, a structure-preserving framework for phenotype-aware molecular representation learning. PhenMol disentangles molecular and cellular representations into shared and private components, enabling phenotype-guided alignment while preserving chemical structures through a dedicated molecular branch. This design integrates cellular phenotype information without disrupting molecular neighborhood organization. Experiments on approximately $3.04 \times 10^{4}$ molecule--cell morphology pairs demonstrate that PhenMol improves molecular property prediction across 270 bioactivity tasks, molecule--phenotype retrieval, and clinical trial outcome prediction. Moreover, ECFP4-based structural analysis shows that PhenMol better preserves molecular neighborhoods and reduces embedding distortion compared with existing multimodal alignment methods. These results highlight the importance of structure-aware constraints in multimodal molecular representation learning and provide an effective approach for integrating cellular phenotypes with chemical knowledge for drug discovery.