Mu-DisCoCat: A Variational Pipeline for Compositional Generalization on Quantum Processors

๐Ÿ“… 2026-10-06
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the scalability bottlenecks of tensor learning and quantum hardware execution challenges in multimodal compositional generalization by proposing a multimodal variational quantum framework. For the first time, this work maps the DisCoCat category to variational quantum circuits and extends it to multimodal scenarios. Adopting a โ€œlearn object representations first, then determine relationsโ€ strategy, the framework leverages Uhlmann fidelity and destructive SWAP tests to achieve image-text compositional understanding. Experiments conducted on both noisy simulators and real IBM quantum processors demonstrate that the model maintains high simulation correlation and surpasses the CLIP baseline in distinguishing unseen compositional pairs. These findings establish a viable pathway for executing compositional generalization on near-term quantum hardware.
๐Ÿ“ Abstract
Achieving compositional concept generalization (CoCoGen), the ability to understand novel situations by recombining learned primitives, remains a fundamental challenge in artificial intelligence. Compositional semantic models such as Compositional Distributional Semantics (DisCoCat) offer solutions by generalising vectors to tensors, but suffer from scaling bottlenecks when learning the tensors. Mapping DisCoCat onto Variational Quantum Circuits (VQCs) resolves this limitation for text, yet the methodology has not been expanded to multimodal situations such as the ones involved in CoCoGen. This paper introduces Mu-DisCoCat: a multimodal variational quantum learning framework for DisCoCat that achieves CoCoGen. The framework first learns stable object representations from single-object image-text pairs, then fixes these and uses them to learn the relations between them in multi-object situations. In classical simulations, the model used Uhlmann state fidelity to compute the overlap between the multimodal circuit representations and achieved higher relational OOD accuracy than the evaluated CLIP baseline. Its deployment was evaluated using the destructive SWAP test across noisy quantum emulators, including a range of IBM fake backends, IQM FakeAphrodite, and the IBM Marrakesh quantum processor. Despite real-world device noise, the hardware-executed models maintained a strong positive correlation with simulated fidelities, reliably distinguishing unseen similar and dissimilar pairs. Our work establishes a framework for executing CoCoGen on VQCs, demonstrating a viable use case for near-term quantum hardware.
Problem

Research questions and friction points this paper is trying to address.

Compositional Generalization
CoCoGen
DisCoCat
Multimodal Learning
Variational Quantum Circuits
Innovation

Methods, ideas, or system contributions that make the work stand out.

Compositional Generalization
Variational Quantum Circuits
Multimodal Learning
DisCoCat
Quantum Processors
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
M
Mina Abbaszadeh
Department of Computer Science, University College London, London, United Kingdom
M
Matilda Karabina Moore
Department of Computer Science, University College London, London, United Kingdom
R
Raem Haq
Department of Computer Science, University College London, London, United Kingdom
Martha Lewis
Martha Lewis
University of Bristol
Artifical IntelligenceCognitive ScienceConceptual SpacesQuantum Theory
Mehrnoosh Sadrzadeh
Mehrnoosh Sadrzadeh
Professor of Computer Science, Royal Academy of Engineering Research Chair,University College London
LogicCategorial GrammarsCompositional Distributional SemanticsMachine Learning