UG-UMRE: Uncertainty-Guided Modality Augmentation and Distributional Calibration for Unified Multimodal Relation Extraction

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing unified multimodal relation extraction approaches, which often suffer from noise propagation due to neglecting inherent aleatoric uncertainty and struggle to align heterogeneous modality distributions. To overcome these challenges, the authors propose the UG-UMRE framework, which first models modality-specific features as Gaussian distributions and employs an uncertainty-aware self-supervised contrastive learning strategy to enhance unimodal representations. Subsequently, it introduces two plug-and-play modules—Joint Aleatoric Uncertainty Alignment (JAUA) and Uncertainty-Driven cross-modal pre-alignment (UDUA)—to construct a shared latent space under a variational information bottleneck constraint, enabling fine-grained cross-modal interaction and distribution calibration. Extensive experiments on UMRE, MORE, and MNRE benchmarks demonstrate that UG-UMRE achieves state-of-the-art performance, confirming its effectiveness and generalizability.
📝 Abstract
Unified Multimodal Relation Extraction (UMRE) aims to identify intra-modal and cross-modal relations between textual entities and visual objects. However, existing UMRE studies still encounter two critical issues: ignoring inherent aleatoric uncertainty causes noise propagation, and deep-seated heterogeneity between distinct modal distributions hinders alignment. To address these issues, we propose the Uncertainty-Guided UMRE Network (UG-UMRE). Specifically, we design an Uncertainty-Driven Unimodal Augmentation (UDUA) module, which models features as Gaussian distributions based on the Variational Information Bottleneck. By incorporating an uncertainty-aware self-supervised contrastive learning mechanism, UDUA effectively filters out noise while maintaining semantic consistency. Furthermore, we introduce the Joint Aleatoric Uncertainty Alignment (JAUA) module as a global semantic pre-calibration mechanism. JAUA leverages probabilistic distribution consistency to construct a shared latent space, eliminating the distributional gap by synchronizing cross-modal statistical properties, thereby laying a robust foundation for fine-grained interaction. Experiments on three benchmark datasets (UMRE, MORE, and MNRE) demonstrate that UG-UMRE achieves state-of-the-art performance. Further analysis validates the pluggable and effective performance of the proposed UDUA and JAUA modules.
Problem

Research questions and friction points this paper is trying to address.

Uncertainty
Modality Heterogeneity
Multimodal Relation Extraction
Noise Propagation
Distribution Alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uncertainty-Guided Learning
Multimodal Relation Extraction
Distributional Calibration
Variational Information Bottleneck
Cross-Modal Alignment
🔎 Similar Papers
No similar papers found.