Score
Design, build, or analyze classifiers and end-to-end pipelines that recognize, retrieve, segment, and evaluate concepts drawn from unconstrained textual vocabularies by mapping text labels into embedding spaces and aligning those embeddings with perceptual (e.g., image-region) representations. This includes vocabulary design and mapping, language- or text-mediated embedding generation, prompt-driven and prompt-free zero-shot classification, detection and instance segmentation, open-vocabulary retrieval and evaluation, and novelty/anomaly (anomnovic/novic) detection workflows, including systems that operate with fixed thresholds or strictly frozen model weights.
This work addresses the challenge faced by non-technical users in formulating accurate natural language category descriptions for open-vocabulary object detection. We propose the first iterative human-in-the-loop feedback mechanism specifically designed for refining such textual class descriptions. Our method integrates text embedding analysis with contrastive example embedding synthesis, enabling users to dynamically define novel categories and iteratively improve description quality *in situ*, without retraining the detector. Evaluated across multiple state-of-the-art open-vocabulary detectors—including GLIP and GroundingDINO—the approach consistently improves detection accuracy (mAP gains of +3.2–5.7), while ensuring interpretability of outputs. Our key contributions are: (1) the first integration of human-AI iterative feedback into textual prompt engineering for open-vocabulary detection; (2) a novel description optimization paradigm grounded in contrastive embedding synthesis; and (3) empirical validation of the mechanism’s cross-model generalizability and robustness.
Existing open-vocabulary semantic segmentation methods rely on manually predefined category names, creating a circular dependency in real-world scenarios: categories must be known *a priori* to perform segmentation. This work introduces the first vocabulary-free semantic segmentation paradigm—requiring no pre-specified lexicon—and enabling fully automatic, pixel-level segmentation via multimodal vision-language models that jointly perceive scene content and generate accurate object descriptions. Our method integrates adaptive text generation, fine-grained cross-modal alignment, and robust text encoding. Crucially, we are the first to reveal the sensitivity of text encoders to generated descriptions and characterize their role in inducing false negatives. Extensive experiments on multiple benchmarks demonstrate significant improvements in segmentation accuracy, validating the strong generalization capability and practical utility of our end-to-end, fully automated pipeline in open-world settings.
Open-vocabulary segmentation significantly lags behind fully supervised methods due to the limitations of vision-language models, which provide only image-level supervision and suffer from semantic ambiguity in natural language. To address this, this work proposes a retrieval-augmented test-time adapter under a few-shot setting that integrates textual prompts with pixel-annotated support images. By leveraging a learnable query-wise cross-modal fusion mechanism, the method dynamically generates lightweight, image-specific classifiers. This approach supports continual expansion of the support set, effectively balancing open-vocabulary generalization with fine-grained segmentation requirements. Extensive experiments demonstrate that it substantially narrows the performance gap between zero-shot and fully supervised segmentation across multiple benchmarks.
This work addresses the weak interpretability of text embedding spaces and their limited structural representation. We propose the Unified Topological Signature (UTS) framework—the first systematic approach to jointly model the topological and geometric structure of embedding spaces. UTS integrates multi-dimensional features, including persistent homology, curvature estimation, and local density, overcoming the redundancy and low discriminability of conventional metrics. By applying clustering analysis and correlation modeling, UTS decodes the mapping between spatial organization and downstream retrieval performance, establishing a quantitative relationship between topological features and document retrievability. Extensive evaluation across multiple state-of-the-art embedding models and benchmark datasets demonstrates that UTS stably predicts inter-model performance differences and ranking effectiveness, exhibiting strong generalization capability and cross-model comparability.
Weak interpretability of large language models (LLMs), unreliability of post-hoc explanation methods, and limitations of concept bottleneck models (CBMs)—including dependence on costly human annotations, restricted representational capacity, and lack of interpretability at the task level—motivate this work. We propose the self-supervised Interpretable Concept Embedding Model (ICEM), the first framework to introduce concept modeling into the textual domain. ICEM leverages the inherent generalization capability of LLMs to autonomously predict concept labels without manual annotation. It enables end-to-end interpretable prediction via concept embeddings and an interpretable decision function, supporting concept intervention, logical attribution, and controllable decoding-path steering. On text classification tasks, ICEM achieves performance comparable to fully supervised CBMs and black-box LLMs, while providing human-understandable, causally grounded explanations. Thus, ICEM unifies interpretability, interactivity, and controllability in a single architecture.
This work addresses the degradation of vision-language prompt alignment in SAM3 under open-vocabulary segmentation due to data and concept drift. To tackle this issue, the authors propose ConceptBank, a dynamic calibration framework that operates without requiring parameter updates. ConceptBank constructs a target-domain-specific concept bank based on statistical characteristics, leveraging class-level visual prototypes as anchors, mining representative support samples, and fusing multiple concepts to jointly correct distribution shifts and restore prompt alignment. Notably, ConceptBank adapts seamlessly to new domains without fine-tuning and significantly enhances the robustness and efficiency of SAM3 in challenging scenarios such as natural scenes and remote sensing imagery, thereby establishing a new benchmark for open-vocabulary segmentation.
This study systematically investigates the intrinsic semantic and syntactic properties of mainstream word embedding methods—such as Word2Vec and GloVe—and their performance disparities across diverse natural language processing tasks. By establishing a unified evaluation framework that integrates publicly available pretrained embeddings with standard benchmark datasets, the work conducts empirical comparisons on canonical tasks including semantic similarity and analogical reasoning. The findings delineate the performance boundaries and optimal application scenarios for each embedding model, offering practitioners reliable guidance for model selection in real-world settings. Furthermore, the analysis deepens the understanding of the inherent limitations of static word representations, highlighting critical constraints in capturing contextual and compositional linguistic phenomena.
This study investigates the relationship between the performance of embedding models and the structural properties of their embedding spaces, with the aim of predicting downstream task effectiveness. Leveraging the MTEB benchmark, the authors evaluate 25 prominent embedding models across five tasks in both English and multilingual settings. They characterize the local and linear structures of embedding spaces using nearest-neighbor overlap and independent component analysis (ICA). The work reveals, for the first time, a remarkably high correlation (up to 0.97) between the degree of local structure preservation in embedding spaces and model performance on downstream tasks. Furthermore, it demonstrates that different tasks exhibit distinct dependencies on local versus linear structural information. These findings indicate that structural characteristics of embedding spaces can effectively predict model performance across diverse tasks, including retrieval, bilingual text mining, pair classification, and summarization.
This work addresses a critical limitation in existing sentence embedding evaluation methods, which rely on downstream classifiers and thus conflate improvements in embedding quality with classifier-induced biases. To overcome this, the authors propose a classifier-free evaluation framework that quantifies how embeddings respond differently to syntactic noise and semantic negation injected into sentences. They introduce the novel “concept separation curve” to visualize a model’s ability to distinguish surface-level perturbations from genuine semantic changes. The approach is validated across multiple languages (English and Dutch), domains, and sentence lengths, demonstrating its effectiveness in providing an interpretable, reproducible, and model-agnostic assessment of conceptual stability in sentence embeddings. This significantly enhances the reliability and transparency of embedding quality evaluation.
This study addresses the challenge of organizing scientific knowledge amid the exponential growth of scholarly literature by proposing an automatic hierarchical classification method based on large language models (LLMs). Leveraging in-context learning (ICL) and prompt chaining, the approach performs three-level categorization—domain, discipline, and topic—within the Open Research Knowledge Graph (ORKG) taxonomy. The first systematic evaluation demonstrates that prompt chaining significantly outperforms conventional ICL, surpassing existing state-of-the-art models particularly at the domain and discipline levels. Although accuracy at the finest-grained topic level remains moderate (approximately 50%), this work validates that off-the-shelf LLMs, without fine-tuning, can effectively support hierarchical semantic organization of scientific texts through carefully engineered prompting strategies.