๐ค AI Summary
This study addresses the cross-modal alignment challenge between powder X-ray diffraction (PXRD) and crystal structures arising from physical information asymmetry. To this end, we propose PXtal, a framework that, for the first time, incorporates physical information asymmetry into the design of multimodal alignment objectives, thereby overcoming the limitations of conventional symmetry assumptions. By jointly leveraging unbalanced optimal transport (UOT) and generalized KullbackโLeibler divergence for supervision, the proposed method achieves efficient representation alignment across these two scientific modalities. Experimental results demonstrate that our approach significantly outperforms existing baselines in zero-shot retrieval tasks and effectively enhances transfer performance on downstream materials science applications.
๐ Abstract
Scientific multimodal learning commonly assumes that paired views are comparably informative. Powder X-ray diffraction (PXRD) makes this mismatch explicit: compressing a three-dimensional crystal structure into a one-dimensional diffraction pattern loses information and makes the pattern harder to connect to the crystal structure that produced it. We introduce PXtal, a framework for learning aligned PXRD and crystal representations under this physically imposed information asymmetry. PXtal uses Unbalanced Optimal Transport (UOT) to adapt the cross-modal coupling and coupling-level generalized Kullback-Leibler (GKL) divergence to supervise the full transport plan. Across six test sets, including four zero-shot transfer sets, PXtal consistently outperforms the baseline models in PXRD-to-crystal candidate retrieval, with the largest gains when PXRD patterns have close but crystallographically distinct nonpaired neighbors, meaning similar input patterns associated with different crystals. The resulting crystal and PXRD encoders transfer more effectively to downstream materials and crystallographic tasks. These results identify information asymmetry as a general design problem in scientific multimodal learning: alignment objectives should reflect what each modality preserves.