build open-vocabulary classifiers

Design, build, or analyze classifiers and end-to-end pipelines that recognize, retrieve, segment, and evaluate concepts drawn from unconstrained textual vocabularies by mapping text labels into embedding spaces and aligning those embeddings with perceptual (e.g., image-region) representations. This includes vocabulary design and mapping, language- or text-mediated embedding generation, prompt-driven and prompt-free zero-shot classification, detection and instance segmentation, open-vocabulary retrieval and evaluation, and novelty/anomaly (anomnovic/novic) detection workflows, including systems that operate with fixed thresholds or strictly frozen model weights.

buildopen-vocabularyclassifiers

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.3
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge faced by non-technical users in formulating accurate natural language category descriptions for open-vocabulary object detection. We propose the first iterative human-in-the-loop feedback mechanism specifically designed for refining such textual class descriptions. Our method integrates text embedding analysis with contrastive example embedding synthesis, enabling users to dynamically define novel categories and iteratively improve description quality *in situ*, without retraining the detector. Evaluated across multiple state-of-the-art open-vocabulary detectors—including GLIP and GroundingDINO—the approach consistently improves detection accuracy (mAP gains of +3.2–5.7), while ensuring interpretability of outputs. Our key contributions are: (1) the first integration of human-AI iterative feedback into textual prompt engineering for open-vocabulary detection; (2) a novel description optimization paradigm grounded in contrastive embedding synthesis; and (3) empirical validation of the mechanism’s cross-model generalizability and robustness.

Enhancing non-technical users' target descriptions with feedbackImproving natural language class descriptions for object detectionOptimizing text embeddings for open-vocabulary detection models

From Open-Vocabulary to Vocabulary-Free Semantic Segmentation

Feb 17, 2025
KR
Klara Reichard
🏛️ BMW Group | Technical University of Munich | University of Padova | Munich Center for Machine Learning | Google Zurich

Existing open-vocabulary semantic segmentation methods rely on manually predefined category names, creating a circular dependency in real-world scenarios: categories must be known *a priori* to perform segmentation. This work introduces the first vocabulary-free semantic segmentation paradigm—requiring no pre-specified lexicon—and enabling fully automatic, pixel-level segmentation via multimodal vision-language models that jointly perceive scene content and generate accurate object descriptions. Our method integrates adaptive text generation, fine-grained cross-modal alignment, and robust text encoding. Crucially, we are the first to reveal the sensitivity of text encoders to generated descriptions and characterize their role in inducing false negatives. Extensive experiments on multiple benchmarks demonstrate significant improvements in segmentation accuracy, validating the strong generalization capability and practical utility of our end-to-end, fully automated pipeline in open-world settings.

Automate object recognition and class naming.Eliminate predefined class vocabularies in segmentation.Enhance vocabulary-free segmentation accuracy.

Open-vocabulary segmentation significantly lags behind fully supervised methods due to the limitations of vision-language models, which provide only image-level supervision and suffer from semantic ambiguity in natural language. To address this, this work proposes a retrieval-augmented test-time adapter under a few-shot setting that integrates textual prompts with pixel-annotated support images. By leveraging a learnable query-wise cross-modal fusion mechanism, the method dynamically generates lightweight, image-specific classifiers. This approach supports continual expansion of the support set, effectively balancing open-vocabulary generalization with fine-grained segmentation requirements. Extensive experiments demonstrate that it substantially narrows the performance gap between zero-shot and fully supervised segmentation across multiple benchmarks.

coarse supervisionfew-shot segmentationopen-vocabulary segmentation

From Topology to Retrieval: Decoding Embedding Spaces with Unified Signatures

Nov 27, 2025
FR
Florian Rottach
🏛️ University of Tübingen | The University of Texas at Austin | Fribourg University

This work addresses the weak interpretability of text embedding spaces and their limited structural representation. We propose the Unified Topological Signature (UTS) framework—the first systematic approach to jointly model the topological and geometric structure of embedding spaces. UTS integrates multi-dimensional features, including persistent homology, curvature estimation, and local density, overcoming the redundancy and low discriminability of conventional metrics. By applying clustering analysis and correlation modeling, UTS decodes the mapping between spatial organization and downstream retrieval performance, establishing a quantitative relationship between topological features and document retrievability. Extensive evaluation across multiple state-of-the-art embedding models and benchmark datasets demonstrates that UTS stably predicts inter-model performance differences and ranking effectiveness, exhibiting strong generalization capability and cross-model comparability.

Analyzing topological and geometric measures of text embedding spacesIntroducing a unified framework to characterize embedding spaces holisticallyLinking topological structure to retrieval performance and model properties

Self-supervised Interpretable Concept-based Models for Text Classification

Jun 20, 2024
FD
Francesco De Santis
🏛️ Politecnico di Torino | Università della Svizzera Italiana | University of Cambridge

Weak interpretability of large language models (LLMs), unreliability of post-hoc explanation methods, and limitations of concept bottleneck models (CBMs)—including dependence on costly human annotations, restricted representational capacity, and lack of interpretability at the task level—motivate this work. We propose the self-supervised Interpretable Concept Embedding Model (ICEM), the first framework to introduce concept modeling into the textual domain. ICEM leverages the inherent generalization capability of LLMs to autonomously predict concept labels without manual annotation. It enables end-to-end interpretable prediction via concept embeddings and an interpretable decision function, supporting concept intervention, logical attribution, and controllable decoding-path steering. On text classification tasks, ICEM achieves performance comparable to fully supervised CBMs and black-box LLMs, while providing human-understandable, causally grounded explanations. Thus, ICEM unifies interpretability, interactivity, and controllability in a single architecture.

Enhancing interpretability of text analysis modelsOvercoming limitations of Concept-Bottleneck Models (CBMs)Reducing dependency on extensive concept annotations

Latest Papers

What's happening recently
View more

This work addresses the degradation of vision-language prompt alignment in SAM3 under open-vocabulary segmentation due to data and concept drift. To tackle this issue, the authors propose ConceptBank, a dynamic calibration framework that operates without requiring parameter updates. ConceptBank constructs a target-domain-specific concept bank based on statistical characteristics, leveraging class-level visual prototypes as anchors, mining representative support samples, and fusing multiple concepts to jointly correct distribution shifts and restore prompt alignment. Notably, ConceptBank adapts seamlessly to new domains without fine-tuning and significantly enhances the robustness and efficiency of SAM3 in challenging scenarios such as natural scenes and remote sensing imagery, thereby establishing a new benchmark for open-vocabulary segmentation.

concept driftdata driftdistribution shift

This study systematically investigates the intrinsic semantic and syntactic properties of mainstream word embedding methods—such as Word2Vec and GloVe—and their performance disparities across diverse natural language processing tasks. By establishing a unified evaluation framework that integrates publicly available pretrained embeddings with standard benchmark datasets, the work conducts empirical comparisons on canonical tasks including semantic similarity and analogical reasoning. The findings delineate the performance boundaries and optimal application scenarios for each embedding model, offering practitioners reliable guidance for model selection in real-world settings. Furthermore, the analysis deepens the understanding of the inherent limitations of static word representations, highlighting critical constraints in capturing contextual and compositional linguistic phenomena.

empirical investigationnatural language processingvector representations

This study investigates the relationship between the performance of embedding models and the structural properties of their embedding spaces, with the aim of predicting downstream task effectiveness. Leveraging the MTEB benchmark, the authors evaluate 25 prominent embedding models across five tasks in both English and multilingual settings. They characterize the local and linear structures of embedding spaces using nearest-neighbor overlap and independent component analysis (ICA). The work reveals, for the first time, a remarkably high correlation (up to 0.97) between the degree of local structure preservation in embedding spaces and model performance on downstream tasks. Furthermore, it demonstrates that different tasks exhibit distinct dependencies on local versus linear structural information. These findings indicate that structural characteristics of embedding spaces can effectively predict model performance across diverse tasks, including retrieval, bilingual text mining, pair classification, and summarization.

benchmark performanceembedding spacesindependent component analysis

This work addresses a critical limitation in existing sentence embedding evaluation methods, which rely on downstream classifiers and thus conflate improvements in embedding quality with classifier-induced biases. To overcome this, the authors propose a classifier-free evaluation framework that quantifies how embeddings respond differently to syntactic noise and semantic negation injected into sentences. They introduce the novel “concept separation curve” to visualize a model’s ability to distinguish surface-level perturbations from genuine semantic changes. The approach is validated across multiple languages (English and Dutch), domains, and sentence lengths, demonstrating its effectiveness in providing an interpretable, reproducible, and model-agnostic assessment of conceptual stability in sentence embeddings. This significantly enhances the reliability and transparency of embedding quality evaluation.

classifier-independentconceptual stabilityevaluation

This study addresses the challenge of organizing scientific knowledge amid the exponential growth of scholarly literature by proposing an automatic hierarchical classification method based on large language models (LLMs). Leveraging in-context learning (ICL) and prompt chaining, the approach performs three-level categorization—domain, discipline, and topic—within the Open Research Knowledge Graph (ORKG) taxonomy. The first systematic evaluation demonstrates that prompt chaining significantly outperforms conventional ICL, surpassing existing state-of-the-art models particularly at the domain and discipline levels. Although accuracy at the finest-grained topic level remains moderate (approximately 50%), this work validates that off-the-shelf LLMs, without fine-tuning, can effectively support hierarchical semantic organization of scientific texts through carefully engineered prompting strategies.

automatic content classificationhierarchical classificationlarge language models

Hot Scholars

FT

Federico Tombari

Google, TU Munich
Computer VisionMachine Learning3D Perception
WN

Wolfgang Nejdl

Professor of Computer Science, Leibniz Universität Hannover, L3S Research Center, Hannover, Germany
Information RetrievalWeb ScienceSocial MediaData Mining
KY

Kailun Yang

Professor. School of Artificial Intelligence and Robotics, Hunan University (HNU); KIT; UAH; ZJU
Computer VisionComputational OpticsIntelligent VehiclesAutonomous Driving
GW

Gaoang Wang

Zhejiang University / University of Illinois Urbana-Champaign Institute
Embodied AgentComputer VisionMachine Learning