Score
Designs and implements perception systems and evaluation pipelines that detect, classify, segment, or retrieve entities from sensor-derived inputs (images, video, or other sensory streams) using arbitrary textual or semantic labels instead of a fixed closed set. This work includes learning and aligning open-ended language and perceptual representations, enabling zero-/few-shot generalization, scalable indexing and retrieval, and mechanisms for open-set/long-tail category handling.
This work addresses the challenge of building video-language understanding systems with human-like perceptual capabilities, enabling synergistic modeling of linguistic and dynamic visual temporal sequences. We systematically survey model architectures, training paradigms, and data construction methodologies in this domain, and introduce— for the first time—a unified, cross-perspective taxonomy that exposes core challenges including multimodal temporal alignment and dataset bias. Leveraging Transformer-based fusion, contrastive/generative pretraining, synthetic data augmentation, and benchmarks such as How2QA and Ego4D, we conduct a comprehensive, reproducible horizontal evaluation of state-of-the-art models under a standardized assessment protocol. Our key contributions are: (1) the first structured analytical framework for joint video-language modeling; (2) clear identification of critical research directions; and (3) a practical, deployable technology roadmap for embodied intelligence.
This work addresses open-vocabulary instance segmentation and online tracking of non-standard objects in dynamic scenes. Methodologically, it introduces the first vision-language model (VLM)-driven, structured-description-guided framework that unifies CLIP/ViT-based semantic understanding, GLIP/OVD-based open-vocabulary detection, and Mask2Former Video for video instance segmentation. Leveraging multimodal prompt engineering and online description generation, it establishes a closed-loop “detection–segmentation–tracking” pipeline. The framework enables zero-shot recognition and attribute-driven trajectory refinement, eliminating reliance on predefined categories or offline training. Evaluated across multiple benchmarks and real-world robotic platforms, it achieves millisecond-level streaming inference, 32.7% mAP for unseen-category segmentation, and 89.2% accuracy in attribute recognition.
This work addresses the neglect of structured relationships among objects within images in open-vocabulary object detection by explicitly modeling semantic and spatial interactions between candidate regions and contextual objects through scene graphs—a first in this domain. The proposed framework integrates a relation-aware attention module with a scene-text alignment branch, jointly leveraging visual relational cues and linguistic semantic knowledge. It further incorporates knowledge distillation and alignment strategies with vision–language models to enhance generalization. Evaluated on the COCO and LVIS benchmarks, the method achieves significant improvements in average precision (AP) for novel categories, outperforming existing open-vocabulary object detection approaches.
Current vision-language models (VLMs) adopt a sequential image processing paradigm, fundamentally misaligning with the human visual system’s parallel dual-path architecture—distinguishing content recognition from spatial perception—and suffer from opaque, uninterpretable internal representations. Method: We propose a decoupling analytical framework that, for the first time, systematically uncovers a two-stage content understanding evolution in VLMs: attribute identification followed by semantic disambiguation; we further establish a geometric theory of positional representation. Leveraging this, we design an instruction-free token compression algorithm and a RoPE scaling technique. Contributions/Results: Through textual image token mapping, visualized decoder analysis, and empirical evaluation, we demonstrate significant improvements in decoding efficiency and spatial reasoning capability. This work establishes an interpretability foundation and actionable design principles for next-generation VLM architectures.
SAM lacks semantic understanding and struggles with open-vocabulary object detection. Method: We propose OS-SAM—the first end-to-end SAM extension enabling open-vocabulary recognition. It introduces a semantic-aware SideFormer module to fuse multimodal features, an open-set RPN that leverages SAM proposals as priors to generate open-vocabulary candidate boxes, and joint localization-classification training via cross-modal alignment and optimization. Contribution/Results: OS-SAM is the first framework to natively adapt SAM for open-vocabulary detection without fine-tuning the image encoder, supporting zero-shot localization and recognition from category names or natural language descriptions. On COCO and LVIS zero-shot detection benchmarks, it significantly outperforms prior state-of-the-art methods, achieving simultaneous improvements in localization accuracy and classification precision—demonstrating a viable pathway for vision foundation models to drive open-vocabulary learning.
This work addresses the challenge of inconsistent semantic representations across multimodal data—such as images, videos, and text—by proposing a language-centric atomic propositional representation framework. The approach transforms observations from any modality into sets of atomic propositions, which are then mapped via a global semantic codebook into a unified, interpretable shared semantic space. This enables compositional expression ranging from fine-grained facts to high-level concepts and facilitates cross-modal reasoning. Experimental results demonstrate that the framework substantially enhances complex multimodal understanding, structured retrieval, and high-quality data curation in autonomous driving and open-world scenarios, offering strong advantages in interpretability, compositionality, and cross-modal alignment.
This work addresses the high computational cost and reliance on large-scale image-text paired data in existing vision-language alignment methods. It proposes, for the first time, repurposing the discarded supervised classification head weights from pretrained vision models as semantic prototypes, enabling efficient zero-shot and few-shot cross-modal alignment without requiring additional paired data. By integrating semantic prototype construction, posterior alignment, and data augmentation, the method consistently enhances the performance of multiple state-of-the-art alignment models across cross-modal retrieval, zero-shot classification, and few-shot classification tasks, significantly outperforming current baselines.
Open-vocabulary segmentation significantly lags behind fully supervised methods due to the limitations of vision-language models, which provide only image-level supervision and suffer from semantic ambiguity in natural language. To address this, this work proposes a retrieval-augmented test-time adapter under a few-shot setting that integrates textual prompts with pixel-annotated support images. By leveraging a learnable query-wise cross-modal fusion mechanism, the method dynamically generates lightweight, image-specific classifiers. This approach supports continual expansion of the support set, effectively balancing open-vocabulary generalization with fine-grained segmentation requirements. Extensive experiments demonstrate that it substantially narrows the performance gap between zero-shot and fully supervised segmentation across multiple benchmarks.
该研究针对视频中开放词汇对象检索问题,提出了一种结构化时空证据图STEG-OVR方法,通过将查询分解为多个槽位来激活相应的检索通道,从而提高检索准确性。
This work investigates the potential of multimodal large language models (MLLMs) to perform purely visual tasks without any training, with a focus on instance-level similarity assessment in large-scale image retrieval. The authors propose a zero-shot reranking method that feeds image pairs into an MLLM and converts its next-token prediction probabilities into similarity scores. Coupled with a memory-efficient indexing mechanism, this approach enables scalable top-k reranking. Notably, it is the first to directly apply MLLMs—without fine-tuning or task-specific architectures—to training-agnostic large-scale image retrieval reranking. The method outperforms specialized rerankers trained on non-native domains across multiple benchmarks and demonstrates superior robustness in challenging scenarios such as cluttered backgrounds, occlusions, and small objects.