🤖 AI Summary
This study addresses the limitation of visual grounding in multimodal large language models, which typically focuses solely on object nouns while neglecting modifiers. To overcome this, we propose OTTER, a method that leverages optimal transport to align generated tokens with visual regions, constructing a compact grounding map. A lightweight supervised probe is then employed to decode modifier-level visual grounding information from frozen model representations. This work provides the first evidence that both modifiers and contextualized nouns encode decodable, discriminative visual cues. Extensive experiments demonstrate that OTTER achieves superior instance-level localization accuracy and robust generalization across free-form generation, cross-dataset transfer, and perturbed scenarios.
📝 Abstract
As Multimodal Large Language Models (MLLMs) can describe increasingly complex visual scenes, token-level grounding becomes crucial. Yet, when an MLLM generates "the yellow banana on the left", established grounding approaches focus on what is in the image ("banana"), overlooking tokens that help describe which instance is meant ("yellow", "left"). In this work, we ask whether frozen MLLM representations contain decodable grounding information about the referred instance across generated tokens, extending to modifiers such as attributes, spatial expressions, and relational/action terms. To address this question, we introduce OTTER, a lightweight supervised probe over frozen MLLM representations that uses Optimal Transport (OT) to align generated tokens with visual regions and produce compact grounding maps. Our results show that (i) instance-discriminative visual information can be decoded from modifier tokens, with the clearest evidence for spatial terms, but (ii) is not confined to them, as contextualized object nouns also carry referential information; (iii) the recovered grounding remains informative under context perturbations, while selected regions remain relevant to generation; and (iv) the learned OT-based grounding extends beyond the controlled setting to free generation and cross-dataset transfer.