MoTIF-X: A Multimodal Tokenized Framework for Interpretable and Extensible Molecular Representation Learning

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitations of insufficient cross-modal interaction and lack of substructure interpretability in multimodal molecular representations by proposing a two-stage pre-training framework anchored on chemical functional groups. Methodologically, it introduces a novel group-centric multimodal alignment mechanism that integrates hierarchical contrastive learning, multimodal masked token modeling, and graph-grounded group extraction to achieve fine-grained fusion of 2D graphs, SMILES strings, and 3D structural information. Beyond enabling substructure-level attribution analysis, the proposed framework achieves state-of-the-art performance on ADMET prediction and drug–target interaction tasks, demonstrating strong generalization capability and scalability for downstream molecular property prediction and drug discovery applications.
πŸ“ Abstract
Molecular representation learning is central to computer-aided drug discovery. Molecular graphs, SMILES strings, and 3D conformations provide complementary structural information, yet many multimodal approaches encode these views independently and align them only at a later stage, limiting fine-grained cross-modal interaction and substructure-level interpretability. To address these limitations, we introduce MoTIF-X, a motif-centered framework that uses graph-grounded chemical motifs as shared anchors for multimodal integration and interpretation. Its first pretraining stage learns motif representations through hierarchical contrastive learning across atomic, motif, and molecular scales. The second stage contextualizes these representations with SMILES and torsion-angle tokens through multimodal masked token modeling. After pretraining on drug-like molecules with multiple conformers, MoTIF-X achieved the lowest mean absolute error on all nine OpenADMET ExpansionRx endpoints and the best overall performance among the evaluated methods. Significance analyses supported its advantage in the vast majority of endpoint-baseline comparisons after multiple-testing correction. Ablation studies supported the complementary contributions of motif-token contextualization, multimodal integration, and two-stage pretraining. Beyond molecular properties, the framework extended to drug-target interaction prediction, achieving the best average classification performance across the evaluated benchmarks and generalizing to an external drug-cold-start dataset without additional fine-tuning. Its motif-centered design also enabled substructure-level interpretation: higher motif attribution scores were associated with larger experimentally measured activity shifts. Together, these findings support MoTIF-X as a transferable and interpretable framework for molecular modeling.
Problem

Research questions and friction points this paper is trying to address.

Molecular representation learning
Multimodal integration
Substructure-level interpretability
Cross-modal interaction
Drug discovery
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal molecular representation
Motif-centered framework
Hierarchical contrastive learning
Masked token modeling
Substructure interpretability
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.