Matching Object or Relation? Tracing Abstract Reasoning Inside VLMs

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the systematic failure of vision-language models (VLMs) on abstract reasoning tasks and the lack of mechanistic explanations for such shortcomings. By integrating the RMTS paradigm with internal mechanism analysis, this work employs representational similarity analysis, causal mediation analysis, and multi-level ablation experiments to investigate the underlying causes and implementation pathways of VLMs' abstract reasoning deficiencies. The findings reveal two competing circuits—one encoding early object features and another capturing late-stage abstract relations—alongside a human-like relational transformation trajectory. Furthermore, four critical leverage points influencing reasoning performance are identified, and the pivotal role of the relational circuit is validated on benchmarks including ARC-AGI.
📝 Abstract
Vision Language Models (VLMs) excel on visual benchmarks but fail systematically on tasks requiring abstract reasoning. Existing benchmarks document this failure but cannot say \emph{why} it happens or which cognitive capability is missing. We close this gap by adopting the Relational Match-to-Sample (RMTS) paradigm from comparative and developmental psychology and pairing it with a mechanistic analysis of the model's internals. On a parametrically controlled stimulus set evaluated across frontier API models (GPT, Claude, Gemini) and three open-source families (Qwen3.5, Gemma-4, InternVL3), we identify four levers that shift VLMs toward the relational match---capability tier, model scale, the number of objects per scene, and the absence of per-object stimulus noise---together producing a developmental-like trajectory that mirrors the human \emph{relational shift}. Opening up the model, a per-layer representational similarity analysis and a causal mediation analysis reveal that VLM abstract reasoning is implemented by two competing circuits: an early circuit that organises images by their surface object features, and a late circuit that organises them by their abstract relation. Extending the analysis to ARC-AGI-1, we find that ablating the relational heads identified on RMTS degrades performance more than ablating random heads, indicating that the relational circuit is recruited beyond our controlled stimuli. We hope this mechanism-level view serves as a step toward understanding how abstract reasoning is implemented in VLMs.
Problem

Research questions and friction points this paper is trying to address.

Vision Language Models
Abstract Reasoning
Relational Match-to-Sample
Mechanistic Analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Abstract Reasoning
Relational Match-to-Sample
Causal Mediation Analysis
Vision Language Models
Mechanistic Interpretability
🔎 Similar Papers
No similar papers found.