🤖 AI Summary
This study addresses the near-random performance of vision-language models (VLMs) in scoring directional spatial relations, identifying a "readout blind spot" where theoretical analysis reveals that symmetric scoring mechanisms cause directional information to mutually cancel out. To overcome this limitation, this work proposes Antisymmetric Displacement Readout (ADR), which operates on frozen encoder features through image-text patch alignment and centroid displacement computation. By incorporating a prior debiasing technique to eliminate interference from textual world knowledge, ADR recovers directional information without requiring additional training. The proposed method significantly outperforms existing scoring paradigms and complex readout approaches at minimal computational cost, achieving highly competitive performance. These results demonstrate that recoverable spatial directional information is inherently preserved within the frozen representations of VLMs.
📝 Abstract
CLIP-like vision-language models remain a cornerstone of multimodal systems, yet their scores stay near chance on directed spatial relations, such as whether one object is left of another. We call this failure readout blindness and analyze, theoretically and empirically, why deployed scores miss the direction: when scoring rules treat the subject and object symmetrically, direction cancels regardless of encoder training. Guided by this analysis, we introduce Antisymmetric Displacement Readout (ADR), which aligns caption words with image patches in the frozen features and scores each relation by the signed displacement between matched object centroids. Notably, ADR succeeds without additional training or learned parameters, thereby demonstrating that directional information remains in the frozen encoder. However, text and world priors can inflate accuracy, so we further introduce prior deflation, which measures the benefit of the image-text pairing as the grounded gain over a null that pairs each item with an unrelated image. Extensive experiments across encoder families show that ADR substantially improves over deployed scores, which remain near chance on most direction-balanced sets even for fine-tuned encoders. Compared with more complex readouts, ADR outperforms the evaluated MLLM likelihood readouts and is competitive with their chat inference at a small fraction of the computation. These results support our claim that directional information can be recovered from frozen features by an appropriate readout. Our implementation and evaluation kit will be publicly available.