🤖 AI Summary
This study addresses the lack of structured inductive biases in multimodal fusion and the deployment challenges of quantum-classical hybrid architectures by proposing the QEMFN framework. This method leverages parameterized quantum entanglement as an inductive bias for cross-modal fusion, constructing a hybrid architecture executable on superconducting quantum devices through angle encoding, cross-modal entanglement circuits, and zero-noise extrapolation. Experimental results demonstrate that QEMFN outperforms multiple classical baselines on the COCO and Flickr30k datasets. Furthermore, quantum centrality evaluation validates the core contribution of the quantum module and its feasibility on real hardware, establishing QEMFN as a valuable and interpretable fusion approach.
📝 Abstract
Multimodal vision-language systems typically fuse image and text embeddings through classical operators such as concatenation, attention, bilinear pooling, or tensor interactions. We propose Quantum Entangled Multimodal Fusion Networks (QEMFN), a hybrid quantum-classical framework that introduces parameterized entanglement as a structured inductive bias for multimodal fusion. Pretrained visual and textual features are projected into compact latent spaces, encoded as angle-parameterized quantum states, processed through intra-modal and paired cross-modal entangling circuits, and measured to produce fused representations for retrieval. Under matched parameter budgets and identical frozen CLIP backbones, QEMFN outperforms classical fusion baselines on COCO-5k and Flickr30k, including multilayer perceptron, tensor fusion, FiLM, cross-attention, compact transformer, and a dequantized paired-topology analogue. An ablation suite isolates the quantum module's contribution from the surrounding classical projections, and quantum-centric analyses report Meyer-Wallach entangling capability, expressibility, gradient variance against barren-plateau bounds, and entropy-performance correlation under controls for training progress alongside an intervention study on the entangling component. QEMFN is executed under shot-based estimation, a noise-modeled fake backend, and a real superconducting device with zero-noise extrapolation. This work does not claim quantum computational advantage; the contribution is the framework together with a controlled empirical and quantum-centric evaluation that positions trainable entanglement as an interpretable, hardware-executable fusion mechanism at scales accessible on contemporary devices.