๐ค AI Summary
This study addresses the scalability bottlenecks of tensor learning and quantum hardware execution challenges in multimodal compositional generalization by proposing a multimodal variational quantum framework. For the first time, this work maps the DisCoCat category to variational quantum circuits and extends it to multimodal scenarios. Adopting a โlearn object representations first, then determine relationsโ strategy, the framework leverages Uhlmann fidelity and destructive SWAP tests to achieve image-text compositional understanding. Experiments conducted on both noisy simulators and real IBM quantum processors demonstrate that the model maintains high simulation correlation and surpasses the CLIP baseline in distinguishing unseen compositional pairs. These findings establish a viable pathway for executing compositional generalization on near-term quantum hardware.
๐ Abstract
Achieving compositional concept generalization (CoCoGen), the ability to understand novel situations by recombining learned primitives, remains a fundamental challenge in artificial intelligence. Compositional semantic models such as Compositional Distributional Semantics (DisCoCat) offer solutions by generalising vectors to tensors, but suffer from scaling bottlenecks when learning the tensors. Mapping DisCoCat onto Variational Quantum Circuits (VQCs) resolves this limitation for text, yet the methodology has not been expanded to multimodal situations such as the ones involved in CoCoGen. This paper introduces Mu-DisCoCat: a multimodal variational quantum learning framework for DisCoCat that achieves CoCoGen. The framework first learns stable object representations from single-object image-text pairs, then fixes these and uses them to learn the relations between them in multi-object situations. In classical simulations, the model used Uhlmann state fidelity to compute the overlap between the multimodal circuit representations and achieved higher relational OOD accuracy than the evaluated CLIP baseline. Its deployment was evaluated using the destructive SWAP test across noisy quantum emulators, including a range of IBM fake backends, IQM FakeAphrodite, and the IBM Marrakesh quantum processor. Despite real-world device noise, the hardware-executed models maintained a strong positive correlation with simulated fidelities, reliably distinguishing unseen similar and dissimilar pairs. Our work establishes a framework for executing CoCoGen on VQCs, demonstrating a viable use case for near-term quantum hardware.