SafeMol: Dual-Modality Safety Alignment for Molecular Multimodal Models

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the safety vulnerabilities and insufficient cross-modal robustness of molecular multimodal models in hazardous molecule scenarios. To this end, it introduces SafeMolBench, a pioneering safety alignment benchmark, and proposes a parameter-efficient multimodal alignment framework. Methodologically, the framework achieves distribution-level cross-modal representation alignment via Maximum Mean Discrepancy, synergistically enhancing model safety through explicit harmful intent modeling and joint text-graph optimization. Experimental results demonstrate that the proposed approach significantly reduces attack success rates while maintaining low over-refusal rates and effectively preserving utility across downstream molecular tasks, thereby achieving an optimal balance among safety, reliability, and practicality.
📝 Abstract
Molecular multimodal models support diverse understanding and generation tasks but may introduce safety vulnerabilities when handling hazardous molecules. In this work, We reveal substantial jailbreak vulnerabilities under both text-only and graph-conditioned settings. Our analysis further shows that safety robustness must hold across input modalities while balancing safety, over-refusal, and utility. To address these challenges, we construct SafeMolBench, a molecular multimodal safety-alignment benchmark with 3702 samples covering 618 unique hazardous molecules and safe molecular tasks, organized into hazardous-harmful, hazardous-allowed, and utility-replay subsets to support unified training and evaluation of safety, over-refusal, and utility. Based on SafeMolBench, we propose SafeMol, a parameter-efficient safety alignment framework that jointly optimizes lightweight modules across text-only and graph-conditioned inputs, uses MMD for distribution-level representation alignment to reduce modality-induced discrepancies, and explicitly models molecular hazardousness and harmful operational intent. Experiments on SafeMolBench show that SafeMol reduces attack success by several tens of percentage points while largely maintaining low over-refusal and preserving molecular-task utility.
Problem

Research questions and friction points this paper is trying to address.

molecular multimodal models
safety alignment
jailbreak vulnerabilities
over-refusal
hazardous molecules
Innovation

Methods, ideas, or system contributions that make the work stand out.

Safety Alignment
Molecular Multimodal Models
Parameter-Efficient Fine-Tuning
Maximum Mean Discrepancy
Jailbreak Vulnerability
🔎 Similar Papers
No similar papers found.
X
Xinmiao Wang
Beihang University
R
Ruijie Wang
Beihang University
M
Menghui Wang
Beihang University
J
Jiawei Chen
Beihang University
H
Haoyue Deng
Beihang University
R
Ran Zhang
China CITIC Bank
Xingxuan Zhang
Xingxuan Zhang
Postdoctoral Research Scientist at Department of Computer Science, Tsinghua University
computer visionOOD GeneralizationDomain GeneralizationOptimization
Xiao Wang
Xiao Wang
Professor, Beihang University
network embeddinggraph neural networksdata miningmachine learning