Diagnosing the Sources of Compositional Failure in Vision-Language Models: A Controlled Analysis

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear failure mechanisms of compositional reasoning in vision-language models (VLMs), where existing methods struggle to disentangle joint reasoning from component recognition costs. We propose COMPASS, a framework introducing a controllable evaluation paradigm to isolate compositional failure factors. By integrating contrastive encoders, explicit compositional reasoning, and fine-grained skill perturbation analysis, it systematically quantifies the compositional integration costs of objects, attributes, and relations. Our findings reveal that intra-load primarily drives degradation, whereas cross-load provides positive contextual support, demonstrating that compositional degradation stems from multifactorial interactions rather than a singular joint reasoning bottleneck. This work offers new perspectives for understanding and enhancing the compositional generalization capabilities of VLMs.
📝 Abstract
Vision-language models (VLMs) often struggle with compositional reasoning tasks, but the reasons for this underperformance remain unclear. A common hypothesis is that models struggle to integrate multiple components, leading to training interventions to improve compositional binding. However, this assumption has never been directly quantified. Existing benchmarks evaluate captions only in their composed form, making it impossible to separate the cost of joint reasoning from the cost of recognizing individual components under increasing load. We introduce COMPASS (COMPositional Analysis of SkillS), a controlled evaluation framework designed to isolate and measure the distinct factors underlying compositional failure. By comparing performance on composed captions with their decomposed counterparts , we directly quantify the cost of compositional integration across 87K image-caption pairs. Across multiple VLMs, this gap is real but partial, accounting for only part of the observed degradation. This motivates a finer-grained investigation into what additional factors govern model behavior. We analyze performance at the level of individual skills: object detection, attribute binding, and relation reasoning, using skill-targeted perturbations across 274K image-caption pairs. We find a consistent skill-specific pattern: each skill degrades primarily with the count of its own primitive type (self-load), while cross-load effects are predominantly positive, suggesting that primitives of different types provide useful grounding context. This pattern holds across standard contrastive encoders, explicitly trained compositional reasoning models, and non-contrastive architectures. These findings show that compositional degradation reflects multiple separable factors that cannot be reduced to joint reasoning alone.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Compositional Reasoning
Compositional Failure
Evaluation Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Compositional Reasoning
Vision-Language Models
Controlled Evaluation Framework
Skill-specific Perturbation
Self-load vs Cross-load
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Mona Gandhi
Mona Gandhi
University of Pennsylvania
Computer VisionNatural Language ProcessingMachine Learning
C
Cenk Merih Olcay
The Ohio State University
K
Kuan-Chieh Lo
The Ohio State University
Santiago Castro
Santiago Castro
Netflix Research
C
Christopher W. Myers
The Ohio State University
S
Srinivasan Parthasarathy
The Ohio State University