Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and Tracking

📅 2025-03-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses open-vocabulary instance segmentation and online tracking of non-standard objects in dynamic scenes. Methodologically, it introduces the first vision-language model (VLM)-driven, structured-description-guided framework that unifies CLIP/ViT-based semantic understanding, GLIP/OVD-based open-vocabulary detection, and Mask2Former Video for video instance segmentation. Leveraging multimodal prompt engineering and online description generation, it establishes a closed-loop “detection–segmentation–tracking” pipeline. The framework enables zero-shot recognition and attribute-driven trajectory refinement, eliminating reliance on predefined categories or offline training. Evaluated across multiple benchmarks and real-world robotic platforms, it achieves millisecond-level streaming inference, 32.7% mAP for unseen-category segmentation, and 89.2% accuracy in attribute recognition.

Technology Category

Computer Vision: Language and VisionMachine Learning: Large Multimodal Models (LMMs)Intelligent Robots: Multimodal Perception & Sensor Fusion

Application Category

Systems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applicationsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved information
📝 Abstract
This paper introduces a novel approach that leverages the capabilities of vision-language models (VLMs) by integrating them with established approaches for open-vocabulary detection (OVD), instance segmentation, and tracking. We utilize VLM-generated structured descriptions to identify visible object instances, collect application-relevant attributes, and inform an open-vocabulary detector to extract corresponding bounding boxes that are passed to a video segmentation model providing precise segmentation masks and tracking capabilities. Once initialized, this model can then directly extract segmentation masks, allowing processing of image streams in real time with minimal computational overhead. Tracks can be updated online as needed by generating new structured descriptions and corresponding open-vocabulary detections. This combines the descriptive power of VLMs with the grounding capability of OVD and the pixel-level understanding and speed of video segmentation. Our evaluation across datasets and robotics platforms demonstrates the broad applicability of this approach, showcasing its ability to extract task-specific attributes from non-standard objects in dynamic environments.
Problem

Research questions and friction points this paper is trying to address.

Integrate vision-language models for open-vocabulary instance segmentation and tracking
Use VLM descriptions to detect objects and generate segmentation masks
Enable real-time processing with minimal computational overhead
Innovation

Methods, ideas, or system contributions that make the work stand out.

Integrates VLMs with open-vocabulary detection and tracking
Uses VLM descriptions for real-time segmentation masks
Combines VLMs, OVD, and video segmentation for dynamic environments
🔎 Similar Papers