Score
Design and build learnable vector representations and small modules that encode semantic concept tokens, together with search, alignment, and optimization procedures to discover, customize, and retrieve instance- or class-specific embeddings at training or test time. Implement interfaces to insert or query these embeddings from downstream models or prompt systems, and develop gradient-based update rules and regularizers to align embeddings with existing token spaces and to reduce spurious or background artifact activations.
This work addresses the challenge in vision-language models (VLMs) of simultaneously achieving discriminative and generative capabilities for custom concepts, along with insufficient composability. We propose a composable custom token learning framework that jointly optimizes text inversion loss and classification loss using only a few images and textual definitions of parent classes, while enforcing regularization via a low-dimensional attribute embedding subspace. We introduce Generative-Aided Information Retrieval (GAIR), a novel paradigm enabling a single custom token to function effectively and coherently across classification, cross-modal retrieval, and text-to-image generation. On DeepFashion2, our method improves mean reciprocal rank (MRR) for text-to-image retrieval by 7%. It further supports interpretable composite query visualization and dynamic inference-time correction, significantly enhancing generalization and controllability for novel concepts.
This study systematically investigates the intrinsic semantic and syntactic properties of mainstream word embedding methods—such as Word2Vec and GloVe—and their performance disparities across diverse natural language processing tasks. By establishing a unified evaluation framework that integrates publicly available pretrained embeddings with standard benchmark datasets, the work conducts empirical comparisons on canonical tasks including semantic similarity and analogical reasoning. The findings delineate the performance boundaries and optimal application scenarios for each embedding model, offering practitioners reliable guidance for model selection in real-world settings. Furthermore, the analysis deepens the understanding of the inherent limitations of static word representations, highlighting critical constraints in capturing contextual and compositional linguistic phenomena.
The evolution of embedding techniques from word vectors to multimodal representations remains fragmented, lacking a unified framework that integrates advances across linguistic, cross-lingual, personalized, and multimodal domains—particularly for embodied multimodal learning in large language models. Method: We systematically survey static and contextual language representations, cross-lingual and personalized modeling, sentence/document embeddings, and multimodal fusion in vision, robotics, and cognitive science. We synthesize recent progress in interpretability, model compression, numerical encoding, and bias mitigation, and propose a novel paradigm emphasizing strong alignment across non-textual modalities and scalable training. Contributions: We construct a comprehensive knowledge graph of end-to-end embedding technologies—from Word2Vec and BERT to GPT, generative topic models, and multimodal alignment/distillation methods—identifying key technical bottlenecks and ethical challenges. This work delivers the first systematic roadmap for multimodal, embodied learning in foundation models.
This work addresses the insufficient robustness of concept representations in large language models (LLMs). We propose the Gaussian Concept Subspace (GCS) framework, which abandons the conventional linear probe assumption that concepts are represented by single vectors, instead modeling concepts as structured subspaces within the model’s latent space. GCS represents each concept as a joint Gaussian distribution parameterized by its mean and covariance matrix, enabling principled subspace confidence estimation and representation-level interventions—e.g., emotion steering. This yields improved semantic stability and conceptual completeness. Extensive experiments across multiple LLM scales and architectures demonstrate that GCS achieves higher faithfulness and plausibility than baseline methods. In emotion-guided generation tasks, GCS simultaneously enhances control precision and text fluency, enabling more controllable and robust concept-level interventions.
This work addresses the low efficiency of key information identification and semantic retrieval in text. We first discover that large language model (LLM) text embeddings naturally align with salient input tokens in the latent space—a phenomenon empirically validated across diverse model architectures, training paradigms, and embedding methods. Leveraging this insight, we propose a principal-component-guided alignment enhancement method that decomposes the embedding geometry and explicitly steers representations toward critical tokens. Building upon this, we design a lightweight sparse retrieval paradigm that retains only ~20% of token embeddings while achieving 80% of dense retrieval performance. Experiments across eight mainstream LLM embedders confirm the robustness of the alignment mechanism. Our findings provide an interpretable foundation for sparse retrieval and instruction-tuned embeddings, and advance the understanding of the intrinsic nature of semantic relevance.
This study investigates the capacity of sentence encoders to represent semantic concepts, revealing that existing models struggle to effectively learn relational and intensional concepts due to mismatches between architecture and supervision signals. Adopting a compositional representation perspective and leveraging a corpus of 3.3 million synonym-definition pairs, the work proposes four guiding principles: fine-tuning with recalibration outperforms expanding the latent space; semantic signals concentrate in the final Transformer layers; hard negative examples enhance discriminability without affecting ranking performance; and supervision efficacy depends on the compositional type of the concept. Through layer-wise pooling ablations, hard negative sampling, and training on large-scale lexical data, the authors construct a new evaluation benchmark—incorporating DBpedia and modifier-annotated noun phrases—and release two novel datasets, offering both theoretical insights and practical resources for research on conceptual representation.
This work addresses a critical limitation in existing sentence embedding evaluation methods, which rely on downstream classifiers and thus conflate improvements in embedding quality with classifier-induced biases. To overcome this, the authors propose a classifier-free evaluation framework that quantifies how embeddings respond differently to syntactic noise and semantic negation injected into sentences. They introduce the novel “concept separation curve” to visualize a model’s ability to distinguish surface-level perturbations from genuine semantic changes. The approach is validated across multiple languages (English and Dutch), domains, and sentence lengths, demonstrating its effectiveness in providing an interpretable, reproducible, and model-agnostic assessment of conceptual stability in sentence embeddings. This significantly enhances the reliability and transparency of embedding quality evaluation.
Current text-to-image generation models struggle to achieve smooth transitions between semantically similar prompts due to substantial differences in token sequences—particularly in wording, ordering, and conceptual positioning—which hinders effective image blending and continuous editing. This work proposes a Token-to-Token Alignment framework that, without modifying the underlying model, employs a two-stage strategy: first aligning the semantic structures of prompts and then aligning their token embedding representations. By reconstructing diverse prompts into a shared structural form, the method reveals that the latent continuous semantic structure within the text embedding space can be effectively leveraged through representation alignment. Consequently, linear interpolation in this aligned space yields coherent semantic transitions, significantly enhancing the quality of image semantic mixing and continuous editing.
This work addresses the suboptimal performance of large language models (LLMs) in plug-and-play text embedding tasks, which stems from their embeddings being overly aligned with high-frequency, low-information tokens in the lexical space, thereby degrading semantic expressiveness. To mitigate this issue, the authors propose EmbedFilter—a training-free, linear filtering method that analyzes the LLM’s unembedding matrix to identify and suppress the subspace implicitly encoding frequent tokens. This approach substantially enhances zero-shot downstream task performance across multiple prominent LLMs. Moreover, EmbedFilter enables significant embedding dimensionality reduction without compromising effectiveness, thereby reducing storage overhead and accelerating retrieval.
This work addresses a key limitation of conventional large language model training, which relies on token-level next-token prediction and consequently struggles to distinguish between semantically equivalent expressions that differ in surface form, leading to a bias toward superficial patterns rather than deep semantic understanding. To overcome this, the authors propose elevating the training objective to the conceptual level by introducing a systematic concept-level supervision signal through a concept mapping framework—e.g., unifying surface variants like “mom” and “mother” under a shared concept such as MOTHER. Their approach integrates a concept alignment loss with a multi-surface aggregation strategy, encouraging the model to prioritize semantic correctness. Experiments demonstrate that the resulting concept-aware models achieve lower perplexity, superior performance across multiple NLP benchmarks, and enhanced robustness in domain transfer scenarios.