Score
Designs, builds, and evaluates segmentation models that explicitly separate and integrate semantic and spatial guidance (or other complementary multimodal cues) by implementing dual-branch or dual-cue architectures. Such work constructs and feeds complementary pixel-wise semantic and spatial cue maps into dedicated segmentation branches, produces cue-specific pixel-wise predictions, and preserves each branch's multimodal reasoning intent before fusing outputs into the final segmentation.
This work investigates whether the understanding and generation branches of unified multimodal models share a transferable semantic space. To this end, the authors propose a cross-branch semantic guidance framework that extracts intervention-based semantic directions from the understanding branch and transfers them to the generation branch for controllable image synthesis. Their experiments reveal, for the first time, a semantic asymmetry between the two branches: the understanding branch encodes object-level semantics, whereas the generation branch relies more heavily on low-level appearance features. The proposed method enables effective semantic transfer from understanding to generation, significantly improving the semantic fidelity of generated images; however, reverse transfer yields limited gains, demonstrating that architectural unification does not inherently guarantee semantic alignment.
R2SM addresses the user-intent-driven mask-type selection problem in text-guided segmentation: given a natural language prompt, the model must determine whether to generate a visible (modal) or complete (amodal) segmentation mask. This work is the first to formulate mask-type decision-making as a language intention understanding task. We introduce the first R2SM benchmark supporting modal/amodal binary classification, constructed by unifying COOA-cls, D2SA, and MUVA datasets, augmented with cross-dataset mask synthesis and fine-grained intention annotation. We propose an end-to-end vision-language framework enabling joint reasoning and mask generation. Experiments demonstrate significant improvements in intention recognition accuracy and mask quality, achieving quantifiable gains in both modal/amodal classification and segmentation precision. Our approach establishes a new paradigm for intention-aware multimodal segmentation.
This work addresses key limitations of large language-multimodal models—namely imprecise object-level localization, difficulty in identity preservation, and low accuracy in region-specific editing—by systematically integrating such models with object-centric visual paradigms for the first time. Focusing on four core tasks—object-level scene understanding, referring expression segmentation, controllable editing, and generation—the study establishes a comprehensive capability pipeline from scene parsing to precise manipulation. The authors introduce an integrated framework grounded in object-centric representations, advancing new directions including instance constancy, spatial control, multi-step interaction consistency, and cross-task unified modeling. They further synthesize pivotal methodologies and evaluation protocols, delineate current capability boundaries, and provide a structured roadmap for reliable assessment and system optimization under distribution shifts.
Do existing promptable segmentation models genuinely understand semantic concepts, or do they merely rely on visually salient yet semantically misleading cues? This work proposes CAFE, a novel benchmark that systematically evaluates conceptual faithfulness from a counterfactual perspective. By constructing attribute-level counterfactual image pairs—where the target region remains unchanged while misleading appearance, context, or material cues are altered—it assesses model robustness against such distractors. Through text-prompt-guided segmentation evaluation and joint analysis of mask accuracy and semantic consistency across 2,146 samples, the study reveals that models frequently produce high-precision masks in response to incorrect prompts, exposing a significant disconnect between their localization capability and true conceptual understanding.
While SAM and SAM 2 excel at segmenting context-agnostic objects (e.g., persons, vehicles), they exhibit significant limitations on context-dependent (CD) concepts—such as visual saliency, camouflaged objects, industrial defects, and medical lesions—primarily due to insufficient global-local semantic co-modeling. Method: We introduce the first comprehensive CD benchmark spanning 11 concept categories across natural, medical, and industrial domains, incorporating 2D/3D images and videos. We propose a unified evaluation framework supporting human annotation, automated metrics, and self-prompted interaction, augmented with prompt robustness testing and SAM 2’s in-context learning analysis. Our method further incorporates multi-granularity prompt generation, cross-modal self-prompting, and context-aware evaluation metrics. Contribution/Results: Experiments reveal fundamental architectural bottlenecks in SAM-series models, delivering the first quantitative analysis of CD segmentation performance and establishing an empirical foundation for designing SAM 3.
This work addresses the unclear mechanisms by which current vision-language models associate spatial relationships with object attributes, particularly the lack of understanding regarding how these models internally process spatial information. Through representational analysis, disentanglement of spatial relations, and enhancement of global visual tokens, the study systematically evaluates the contribution of individual components to spatial reasoning. It reveals, for the first time, that the visual encoder plays a dominant role in spatial reasoning: its output encodes global spatial signals distributed broadly across all image tokens—including background regions—rather than being confined to object-centric areas. Leveraging this insight, augmenting the visual encoder’s global spatial representations substantially improves spatial reasoning performance on natural images, challenging the conventional paradigm that focuses exclusively on object regions.
This work addresses the limitations of existing language-guided segmentation methods, which rely heavily on large-scale training and lack explicit visual-spatial reasoning capabilities, thereby struggling to accurately segment arbitrarily described targets in zero-shot settings. To overcome these challenges, the authors propose Seg-Agent—a training-free framework that enables explicit multimodal reasoning through an iterative loop of generation, selection, and refinement. By integrating Set-of-Mark visual prompts, Seg-Agent orchestrates collaborative spatial reasoning between a multimodal large language model and a foundation segmentation model such as SAM. This approach introduces the first explicit multimodal chain-of-reasoning mechanism, transcending the constraints of purely textual inference. Without any parameter updates, Seg-Agent achieves performance comparable to state-of-the-art trained methods and establishes Various-LangSeg, a new benchmark encompassing semantic, generic object, and complex reasoning segmentation tasks.
This study addresses the instability of multimodal large language models (MLLMs) in pixel-level tasks such as semantic segmentation and investigates their poorly understood spatial reasoning mechanisms. Through systematic layer-wise linear probing and attention intervention analyses, the authors evaluate the representational capabilities across the visual encoder, adapter, and large language model (LLM) stages. They uncover a previously unknown “degradation–recovery” mechanism: segmentation performance degrades within the adapter but is progressively restored in the LLM via attention mechanisms. Furthermore, the work demonstrates that correctly classified image tokens can guide neighboring misclassified tokens toward correction through bidirectional attention, effectively mitigating the limitations imposed by causal attention. These findings provide crucial mechanistic insights and architectural guidance for designing MLLMs with robust segmentation capabilities.
Multimodal language models struggle to effectively leverage pixel-level dense representations from vision tools—such as depth or optical flow—resulting in limited perceptual capabilities and overreliance on linguistic priors. This work proposes Perception Programs (P²), a training-free, model-agnostic approach that, for the first time, translates vision tool outputs into compact, structured, and language-native natural language summaries via procedural rules, enabling direct parsing and reasoning by language models. Departing from the conventional paradigm of feeding raw pixel features, P² achieves an average 22% improvement across six perception tasks in the BLINK benchmark. Notably, GPT-5 Mini attains 86.47% accuracy in multi-view reasoning and 81.45% in relative depth estimation, while smaller models also gain absolute improvements of 15–40%, establishing new state-of-the-art results.
This study investigates the impact of external information—such as spatial cues, commonsense knowledge, and chain-of-thought prompts—on visual spatial reasoning (VSR) performance. Through hypothesis-driven controlled experiments on two public benchmarks, the authors systematically evaluate three categories of vision-language models. Their findings reveal that a single, precise spatial cue consistently outperforms multi-context fusion; weakly relevant or excessive commonsense knowledge degrades performance; and chain-of-thought prompting is beneficial only when spatial localization is sufficiently accurate. This work is the first to demonstrate that “more information is not always better” in VSR, advocating instead for the selective injection of task-aligned signals and clarifying the conditions under which spatial localization and chain-of-thought reasoning synergistically enhance performance.