PRISM: A Category-Theoretic Framework for Measuring and Refining Multimodal Analogies

πŸ“… 2026-10-01
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the lack of interpretable metrics and the difficulty of verifying relational structure preservation in existing multimodal analogy generation by proposing the PRISM framework. Grounded in category theory, PRISM formalizes analogies as explicit relational mappings and leverages vision-language models (VLMs) for cross-modal instantiation. It introduces a novel "pullback score" to quantify relational alignment, which subsequently drives an iterative feedback refinement mechanism to optimize visual metaphor generation. Evaluated on the AnaloBench benchmark, analogy selection based solely on the pullback score achieves 82.5% accuracy, while human evaluations confirm significant improvements in refined output quality. This work establishes a structured and interpretable paradigm for multimodal analogical reasoning.
πŸ“ Abstract
Analogical reasoning involves identifying and preserving relational structures across domains. However, existing approaches to AI-driven multimodal analogy generation lack an interpretable measure of whether this structure is understood and maintained in the generated output. We address this gap with Pullback Refinement via Interpretable Structural Mapping (PRISM), a modality-agnostic framework for measuring and improving relational alignment in multimodal analogies, evaluated on visual metaphor generation. PRISM represents analogies as explicit relational mappings grounded in category theory and uses VLMs to instantiate these structures across modalities. Its first component, the pullback score, quantifies relational alignment from the resulting graph representation. On the AnaloBench benchmark, selecting the correct analogy purely by pullback score achieves 82.5% accuracy, demonstrating that the score captures meaningful relational information. PRISM's second component is an iterative refinement loop that uses the pullback score as an in-context feedback signal to iteratively revise the generated image towards greater relational depth. VLM-as-a-judge and human evaluations show that PRISM consistently improves metaphor consistency and analogy appropriateness over zero- shot generation, with human participants preferring the refined output in 57.65% of pairwise comparisons. However, a qualitative analysis reveals that refinement can favour visually crowded compositions rather than genuinely deeper relational correspondences.
Problem

Research questions and friction points this paper is trying to address.

multimodal analogy
relational structure
interpretable measure
analogical reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Category Theory
Multimodal Analogies
Pullback Score
Iterative Refinement
Vision-Language Models
πŸ”Ž Similar Papers
No similar papers found.
M
Mirella Zeisler
Delft University of Technology
O
Ojas Shirekar
Delft University of Technology
M
Mircea Licǎ
Delft University of Technology
Chirag Raman
Chirag Raman
Delft University of Technology
Multimodal Machine LearningComputer VisionHuman-Computer Interaction