Who Is Left of Whom? Tracing Spatial Evidence and Role Binding in Relative-Position Reasoning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inconsistency in vision-language models during relative spatial reasoning caused by object role swapping, revealing a hierarchical mechanism underlying spatial evidence tracking and role binding. Methodologically, it localizes critical representation layers via activation patching to elucidate the staged progression from source expressions to query roles, establishes causal pathways, and identifies stable role directions. Interventions are then implemented through directional manipulation combined with direction-guidance techniques on synthetic scenes. Results demonstrate that this approach significantly improves both reasoning accuracy and pairwise consistency on natural image benchmarks without retraining, offering an effective training-free intervention paradigm for enhancing the robustness of spatial reasoning in foundation models.
📝 Abstract
High instance-level accuracy can mask inconsistencies in spatial reasoning when objects exchange positions or their roles are reversed in the query. The internal representations supporting relative-position reasoning remain poorly understood. We investigate two complementary components of this process: tracking object locations in the input and representing their query roles. Across three VLMs with visual or textual inputs and their language-model backbones, activation patching reveals a staged progression from early-layer source representations through intermediate-layer query-object representations to late-layer answer states. Targeted interventions further establish causal links along this progression: manipulating source-side representations shifts location information at query-object mentions and ultimately alters relation predictions. Beyond object-location information, we also identify a stable query-side direction associated with the roles of the two objects in the comparison. Steering along directions estimated on synthetic scenes generalizes to natural-image benchmarks, improving accuracy and both forms of paired consistency in most settings without retraining. Our findings reveal complementary components of relational reasoning across visual and textual settings and show how targeted interventions can improve the consistency of models'behavior.
Problem

Research questions and friction points this paper is trying to address.

relative-position reasoning
spatial reasoning
role binding
visual language models
consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Activation Patching
Relative-Position Reasoning
Role Binding
Representation Steering
Causal Intervention
🔎 Similar Papers
No similar papers found.