develop open-vocabulary perception

Designs and implements perception systems and evaluation pipelines that detect, classify, segment, or retrieve entities from sensor-derived inputs (images, video, or other sensory streams) using arbitrary textual or semantic labels instead of a fixed closed set. This work includes learning and aligning open-ended language and perceptual representations, enabling zero-/few-shot generalization, scalable indexing and retrieval, and mechanisms for open-set/long-tail category handling.

developopen-vocabularyperception

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.34
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives

Jun 09, 2024
TN
Thong Nguyen
🏛️ National University of Singapore | Nanyang Technological University

This work addresses the challenge of building video-language understanding systems with human-like perceptual capabilities, enabling synergistic modeling of linguistic and dynamic visual temporal sequences. We systematically survey model architectures, training paradigms, and data construction methodologies in this domain, and introduce— for the first time—a unified, cross-perspective taxonomy that exposes core challenges including multimodal temporal alignment and dataset bias. Leveraging Transformer-based fusion, contrastive/generative pretraining, synthetic data augmentation, and benchmarks such as How2QA and Ego4D, we conduct a comprehensive, reproducible horizontal evaluation of state-of-the-art models under a standardized assessment protocol. Our key contributions are: (1) the first structured analytical framework for joint video-language modeling; (2) clear identification of critical research directions; and (3) a practical, deployable technology roadmap for embodied intelligence.

Analyze challenges in model training for video-language tasksCompare performance and future research directionsSurvey video-language understanding systems' model architectures

This work addresses open-vocabulary instance segmentation and online tracking of non-standard objects in dynamic scenes. Methodologically, it introduces the first vision-language model (VLM)-driven, structured-description-guided framework that unifies CLIP/ViT-based semantic understanding, GLIP/OVD-based open-vocabulary detection, and Mask2Former Video for video instance segmentation. Leveraging multimodal prompt engineering and online description generation, it establishes a closed-loop “detection–segmentation–tracking” pipeline. The framework enables zero-shot recognition and attribute-driven trajectory refinement, eliminating reliance on predefined categories or offline training. Evaluated across multiple benchmarks and real-world robotic platforms, it achieves millisecond-level streaming inference, 32.7% mAP for unseen-category segmentation, and 89.2% accuracy in attribute recognition.

Enable real-time processing with minimal computational overheadIntegrate vision-language models for open-vocabulary instance segmentation and trackingUse VLM descriptions to detect objects and generate segmentation masks

This work addresses the neglect of structured relationships among objects within images in open-vocabulary object detection by explicitly modeling semantic and spatial interactions between candidate regions and contextual objects through scene graphs—a first in this domain. The proposed framework integrates a relation-aware attention module with a scene-text alignment branch, jointly leveraging visual relational cues and linguistic semantic knowledge. It further incorporates knowledge distillation and alignment strategies with vision–language models to enhance generalization. Evaluated on the COCO and LVIS benchmarks, the method achieves significant improvements in average precision (AP) for novel categories, outperforming existing open-vocabulary object detection approaches.

novel categoriesopen-vocabulary object detectionscene graphs

Reading Images Like Texts: Sequential Image Understanding in Vision-Language Models

Sep 23, 2025
YL
Yueyan Li
🏛️ Beijing University of Posts and Telecommunications

Current vision-language models (VLMs) adopt a sequential image processing paradigm, fundamentally misaligning with the human visual system’s parallel dual-path architecture—distinguishing content recognition from spatial perception—and suffer from opaque, uninterpretable internal representations. Method: We propose a decoupling analytical framework that, for the first time, systematically uncovers a two-stage content understanding evolution in VLMs: attribute identification followed by semantic disambiguation; we further establish a geometric theory of positional representation. Leveraging this, we design an instruction-free token compression algorithm and a RoPE scaling technique. Contributions/Results: Through textual image token mapping, visualized decoder analysis, and empirical evaluation, we demonstrate significant improvements in decoding efficiency and spatial reasoning capability. This work establishes an interpretability foundation and actionable design principles for next-generation VLM architectures.

Analyzes object recognition and spatial perception in Vision-Language ModelsIdentifies a two-stage process from attributes to semantic disambiguationProposes methods to improve decoding efficiency and spatial reasoning

Boosting Segment Anything Model Towards Open-Vocabulary Learning

Dec 06, 2023
XH
Xumeng Han
🏛️ University of Chinese Academy of Sciences | Huawei Cloud

SAM lacks semantic understanding and struggles with open-vocabulary object detection. Method: We propose OS-SAM—the first end-to-end SAM extension enabling open-vocabulary recognition. It introduces a semantic-aware SideFormer module to fuse multimodal features, an open-set RPN that leverages SAM proposals as priors to generate open-vocabulary candidate boxes, and joint localization-classification training via cross-modal alignment and optimization. Contribution/Results: OS-SAM is the first framework to natively adapt SAM for open-vocabulary detection without fine-tuning the image encoder, supporting zero-shot localization and recognition from category names or natural language descriptions. On COCO and LVIS zero-shot detection benchmarks, it significantly outperforms prior state-of-the-art methods, achieving simultaneous improvements in localization accuracy and classification precision—demonstrating a viable pathway for vision foundation models to drive open-vocabulary learning.

Enhance SAM for open-vocabulary object detectionImprove localization and classification in zero-shot scenariosIntegrate semantic information into SAM features

Latest Papers

What's happening recently
View more

This work addresses the challenge of inconsistent semantic representations across multimodal data—such as images, videos, and text—by proposing a language-centric atomic propositional representation framework. The approach transforms observations from any modality into sets of atomic propositions, which are then mapped via a global semantic codebook into a unified, interpretable shared semantic space. This enables compositional expression ranging from fine-grained facts to high-level concepts and facilitates cross-modal reasoning. Experimental results demonstrate that the framework substantially enhances complex multimodal understanding, structured retrieval, and high-quality data curation in autonomous driving and open-world scenarios, offering strong advantages in interpretability, compositionality, and cross-modal alignment.

cross-modal understandinginterpretable AImultimodal representation

This work addresses the high computational cost and reliance on large-scale image-text paired data in existing vision-language alignment methods. It proposes, for the first time, repurposing the discarded supervised classification head weights from pretrained vision models as semantic prototypes, enabling efficient zero-shot and few-shot cross-modal alignment without requiring additional paired data. By integrating semantic prototype construction, posterior alignment, and data augmentation, the method consistently enhances the performance of multiple state-of-the-art alignment models across cross-modal retrieval, zero-shot classification, and few-shot classification tasks, significantly outperforming current baselines.

Cross-modal RetrievalSemantic PrototypesVision-Language Alignment

Open-vocabulary segmentation significantly lags behind fully supervised methods due to the limitations of vision-language models, which provide only image-level supervision and suffer from semantic ambiguity in natural language. To address this, this work proposes a retrieval-augmented test-time adapter under a few-shot setting that integrates textual prompts with pixel-annotated support images. By leveraging a learnable query-wise cross-modal fusion mechanism, the method dynamically generates lightweight, image-specific classifiers. This approach supports continual expansion of the support set, effectively balancing open-vocabulary generalization with fine-grained segmentation requirements. Extensive experiments demonstrate that it substantially narrows the performance gap between zero-shot and fully supervised segmentation across multiple benchmarks.

coarse supervisionfew-shot segmentationopen-vocabulary segmentation

This work investigates the potential of multimodal large language models (MLLMs) to perform purely visual tasks without any training, with a focus on instance-level similarity assessment in large-scale image retrieval. The authors propose a zero-shot reranking method that feeds image pairs into an MLLM and converts its next-token prediction probabilities into similarity scores. Coupled with a memory-efficient indexing mechanism, this approach enables scalable top-k reranking. Notably, it is the first to directly apply MLLMs—without fine-tuning or task-specific architectures—to training-agnostic large-scale image retrieval reranking. The method outperforms specialized rerankers trained on non-native domains across multiple benchmarks and demonstrates superior robustness in challenging scenarios such as cluttered backgrounds, occlusions, and small objects.

Image RetrievalInstance-level SimilarityLarge-scale Vision Tasks

Hot Scholars

PZ

Pengsong Zhang

Ph.D. Candidate, University of Toronto
RoboticsAgentComputer visionReinforcement learning
XW

Xueqian Wang

Tsinghua University
Information FusionTarget DetectionRadar ImagingImage Processing
EM

Eduardo Montijano

Universidad de Zaragoza, Spain
Roboticscomputer visiondistributed systems