Quantum Entangled Multimodal Fusion Networks (QEMFN): Resource-Aware Hybrid Vision-Language Fusion via Trainable Entanglement

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of structured inductive biases in multimodal fusion and the deployment challenges of quantum-classical hybrid architectures by proposing the QEMFN framework. This method leverages parameterized quantum entanglement as an inductive bias for cross-modal fusion, constructing a hybrid architecture executable on superconducting quantum devices through angle encoding, cross-modal entanglement circuits, and zero-noise extrapolation. Experimental results demonstrate that QEMFN outperforms multiple classical baselines on the COCO and Flickr30k datasets. Furthermore, quantum centrality evaluation validates the core contribution of the quantum module and its feasibility on real hardware, establishing QEMFN as a valuable and interpretable fusion approach.
📝 Abstract
Multimodal vision-language systems typically fuse image and text embeddings through classical operators such as concatenation, attention, bilinear pooling, or tensor interactions. We propose Quantum Entangled Multimodal Fusion Networks (QEMFN), a hybrid quantum-classical framework that introduces parameterized entanglement as a structured inductive bias for multimodal fusion. Pretrained visual and textual features are projected into compact latent spaces, encoded as angle-parameterized quantum states, processed through intra-modal and paired cross-modal entangling circuits, and measured to produce fused representations for retrieval. Under matched parameter budgets and identical frozen CLIP backbones, QEMFN outperforms classical fusion baselines on COCO-5k and Flickr30k, including multilayer perceptron, tensor fusion, FiLM, cross-attention, compact transformer, and a dequantized paired-topology analogue. An ablation suite isolates the quantum module's contribution from the surrounding classical projections, and quantum-centric analyses report Meyer-Wallach entangling capability, expressibility, gradient variance against barren-plateau bounds, and entropy-performance correlation under controls for training progress alongside an intervention study on the entangling component. QEMFN is executed under shot-based estimation, a noise-modeled fake backend, and a real superconducting device with zero-noise extrapolation. This work does not claim quantum computational advantage; the contribution is the framework together with a controlled empirical and quantum-centric evaluation that positions trainable entanglement as an interpretable, hardware-executable fusion mechanism at scales accessible on contemporary devices.
Problem

Research questions and friction points this paper is trying to address.

multimodal fusion
vision-language systems
quantum entanglement
resource-aware
hybrid quantum-classical
Innovation

Methods, ideas, or system contributions that make the work stand out.

Quantum Entangled Multimodal Fusion
Hybrid Quantum-Classical Framework
Trainable Entanglement
Vision-Language Retrieval
Zero-Noise Extrapolation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Srikar Alla
Department of Electrical Engineering and Computer Science, University of Missouri, Columbia, MO 65211, USA
A
Ali Shiri Sichani
Department of Electrical Engineering and Computer Science, University of Missouri, Columbia, MO 65211, USA
Chi-Ren Shyu
Chi-Ren Shyu
Director of MU Institute for Data Science and Informatics, Associate Dean of Engineering
Biomedical InformaticsGeoinformaticsQuantum Computing