Score
Designs and builds models, pipelines, or precomputed representations that encode textual inputs (tokens, phrases, sentences, or node texts) into a single shared semantic vector space. This includes creating or optimizing embedding models, producing frozen/general encoders, and ensuring cross-dataset textual consistency so the embeddings can serve as unified input features for downstream models.
Existing general-purpose text embedding (GPTE) methods exhibit fragmented understanding of pre-trained language models’ (PLMs) roles and rely on narrow optimization objectives. Method: We propose a multi-level framework categorizing PLMs’ roles across architecture design, representation enhancement, training paradigms, and data construction; extend the scope to emerging challenges—including safety, bias mitigation, and cognitive scalability—and integrate PLM-driven dense vector generation, contrastive learning, and large-scale pairwise supervision to support multilingual, multimodal, and code embedding. Contribution/Results: This work delivers the first structured technical survey and developmental roadmap for GPTE, systematically clarifying PLMs’ evolving functions. It significantly improves generalization and robustness across downstream tasks—including retrieval, classification, and clustering—while unifying disparate methodological advances under a coherent conceptual framework.
This work addresses the joint optimization of large language models (LLMs) and text embedding techniques to enhance efficiency and robustness in semantic matching, clustering, and information retrieval. We propose the first unified taxonomy centered on the *interaction patterns* between LLMs and embeddings—categorizing approaches into three paradigms: LLM-augmented embeddings, LLM-as-embedder, and LLM-understanding-embeddings—thereby transcending conventional task-centric taxonomies. By integrating supervised/unsupervised embedding learning, instruction tuning, prompt engineering, representation space analysis, and interpretability methods, we construct a structured knowledge graph encompassing over 100 studies. Our framework precisely delineates capability boundaries and application scopes for each paradigm, identifies persistent limitations inherited from pre-trained language models (PLMs) and novel challenges introduced by LLMs, and provides a theoretically grounded, empirically informed roadmap for future advancement.
This study investigates whether latent representations from heterogeneous text embedding models can be transferred via simple transformations to enable direct AI-to-AI communication without decoding into human-readable text. For the first time, we systematically evaluate the effectiveness and limitations of linear mappings as lightweight translators across nine diverse models varying in architecture, pooling strategy, and training objective, using real-world textual data. Through comprehensive metrics—including Centered Kernel Alignment (CKA) similarity, downstream task transfer performance, fidelity, and retrieval accuracy—we find that simple transformations succeed only between partially compatible model pairs and largely fail otherwise. These results indicate that semantic transfer across heterogeneous embedding spaces cannot be universally achieved through alignment alone, as compatibility is jointly constrained by architectural design, training objectives, pooling mechanisms, and data distribution.
The evolution of embedding techniques from word vectors to multimodal representations remains fragmented, lacking a unified framework that integrates advances across linguistic, cross-lingual, personalized, and multimodal domains—particularly for embodied multimodal learning in large language models. Method: We systematically survey static and contextual language representations, cross-lingual and personalized modeling, sentence/document embeddings, and multimodal fusion in vision, robotics, and cognitive science. We synthesize recent progress in interpretability, model compression, numerical encoding, and bias mitigation, and propose a novel paradigm emphasizing strong alignment across non-textual modalities and scalable training. Contributions: We construct a comprehensive knowledge graph of end-to-end embedding technologies—from Word2Vec and BERT to GPT, generative topic models, and multimodal alignment/distillation methods—identifying key technical bottlenecks and ethical challenges. This work delivers the first systematic roadmap for multimodal, embodied learning in foundation models.
This work investigates how large language models encode non-factual literary style—particularly authorial style—in deep prompt embeddings, beyond semantic or factual content representation. Methodologically, it leverages Transformer-based architectures and employs geometric analysis of embedding spaces, cross-text style clustering, and visualization to systematically characterize the distributional properties of short texts in high-dimensional latent space. The study reveals, for the first time, that deep embeddings of texts by the same author exhibit strong aggregation and entanglement, while those from different authors are markedly separated; moreover, the geometric structure of these embeddings stably encodes abstract stylistic features. These findings demonstrate that prompt embeddings serve not merely as compressed semantic representations but also as compact, structured encodings of stylistic information. Consequently, this work establishes a novel, interpretable, and computationally tractable paradigm for author attribution, stylometry, and related tasks.
This study challenges the implicit assumption that geometric proximity (e.g., cosine similarity) in sentence embedding spaces reflects semantic or functional similarity, asking whether such geometric properties can predict relative performance on downstream language tasks. Method: Within a unified Transformer framework, we systematically compare three embedding strategies—mean-pooled token, [CLS] token, and randomly selected token embeddings—across multiple NLP tasks. We conduct rigorous distance–performance correlation analysis to assess how well cosine similarity predicts task accuracy. Contribution/Results: We find that cosine similarity captures only shallow, surface-level lexical commonalities and fails to reliably predict downstream performance. Crucially, task-relevant semantic similarity is encoded via dimensionally weighted combinations rather than isotropic geometric proximity; thus, embeddings with large geometric distances in high-dimensional space may still encode highly similar task-specific semantics. This work provides the first empirical evidence of a substantial decoupling between the geometric structure of sentence embeddings and their functional utility, establishing a new paradigm for embedding evaluation and design.
This work addresses the inefficiency of sentence embedding extraction from pretrained Transformers (e.g., BERT). We systematically investigate and enhance three key strategies: token aggregation, representation post-processing, and external-knowledge-guided fine-tuning. Specifically, we propose novel representation shaping techniques—including weighted aggregation of multi-layer hidden states, normalized contrastive fine-tuning, and Wikidata-augmented supervision—achieving substantial improvements in semantic expressiveness of static or randomly initialized embeddings, without introducing additional parameters or inference overhead. Our approach outperforms strong baselines across 8 semantic textual similarity, 6 short-text clustering, and 12 classification tasks. Notably, optimized random embeddings achieve over 120% improvement on STS-B, approaching native BERT performance. Empirical results validate the effectiveness and cross-model generalizability of lightweight representation shaping for universal sentence embedding learning.
This study systematically investigates the intrinsic semantic and syntactic properties of mainstream word embedding methods—such as Word2Vec and GloVe—and their performance disparities across diverse natural language processing tasks. By establishing a unified evaluation framework that integrates publicly available pretrained embeddings with standard benchmark datasets, the work conducts empirical comparisons on canonical tasks including semantic similarity and analogical reasoning. The findings delineate the performance boundaries and optimal application scenarios for each embedding model, offering practitioners reliable guidance for model selection in real-world settings. Furthermore, the analysis deepens the understanding of the inherent limitations of static word representations, highlighting critical constraints in capturing contextual and compositional linguistic phenomena.
Large language models (LLMs) struggle to produce high-quality holistic text representations due to their pretraining objective—autoregressive token-level prediction—which inherently prioritizes local lexical coherence over global semantic structure. To address this, we propose *context compression*, a novel unsupervised pretraining task wherein the model encodes long input contexts into a compact sequence of memory tokens and reconstructs the original sequence from them. This paradigm shifts focus from token-level modeling to holistic representation learning and explicitly enforces global semantic consistency via contrastive learning. Based on this objective, we introduce LLM2Comp—a lightweight, efficient encoder derived from frozen LLM backbones. Empirical evaluation shows that LLM2Comp significantly outperforms state-of-the-art LLM-based text encoders (e.g., Instructor, BGE) on downstream tasks including text classification and semantic retrieval, while requiring only 20–33% of their training data. It achieves superior sample efficiency, stronger generalization across domains, and reduced inference latency.
This work addresses the weak interpretability of text embedding spaces and their limited structural representation. We propose the Unified Topological Signature (UTS) framework—the first systematic approach to jointly model the topological and geometric structure of embedding spaces. UTS integrates multi-dimensional features, including persistent homology, curvature estimation, and local density, overcoming the redundancy and low discriminability of conventional metrics. By applying clustering analysis and correlation modeling, UTS decodes the mapping between spatial organization and downstream retrieval performance, establishing a quantitative relationship between topological features and document retrievability. Extensive evaluation across multiple state-of-the-art embedding models and benchmark datasets demonstrates that UTS stably predicts inter-model performance differences and ranking effectiveness, exhibiting strong generalization capability and cross-model comparability.
This study investigates the reliance of text-to-image generation models on linguistic information encoded in their text encoders. To this end, the authors propose a decontextualized “bag-of-words with positional tags” embedding approach that preserves only lexical identity and word order while discarding rich semantic context, and employ it to guide a diffusion Transformer in image synthesis. Experimental results demonstrate that this simplified embedding achieves comparable performance to full contextual embeddings in both visual quality and text-image alignment. These findings reveal, for the first time, that state-of-the-art models do not critically depend on deep linguistic structures from the text encoder; instead, semantic integration is largely accomplished by the image generator itself, thereby challenging conventional assumptions about the role of text encoders in generative modeling.
This work addresses the suboptimal performance of large language models (LLMs) in plug-and-play text embedding tasks, which stems from their embeddings being overly aligned with high-frequency, low-information tokens in the lexical space, thereby degrading semantic expressiveness. To mitigate this issue, the authors propose EmbedFilter—a training-free, linear filtering method that analyzes the LLM’s unembedding matrix to identify and suppress the subspace implicitly encoding frequent tokens. This approach substantially enhances zero-shot downstream task performance across multiple prominent LLMs. Moreover, EmbedFilter enables significant embedding dimensionality reduction without compromising effectiveness, thereby reducing storage overhead and accelerating retrieval.