Score
Fine-tune Sentence‑BERT / Sentence‑Transformers encoders to produce dense sentence embeddings whose vector geometry (cosine distance or dot product) encodes semantic similarity, using objectives such as cosine‑similarity regression, contrastive, or triplet losses. Design the finetuning pipeline — datasets and sampling, loss formulation, training and evaluation (e.g., retrieval recall@k, clustering, ranking) — and deliver models and embeddings for embedding‑based retrieval, clustering, or similarity scoring.
This study identifies an anomalous performance degradation—where fine-tuning strong pretrained Transformer models (e.g., BERT, DeBERTa) on the MS MARCO passage ranking task reduces MRR@10 below the base model’s score (0.3026)—across all five fine-tuning strategies, including full-parameter tuning and parameter-efficient methods like LoRA. Using UMAP-based embedding visualization, training dynamics analysis, and efficiency evaluation, we demonstrate that fine-tuning disrupts the optimal semantic embedding structure acquired during large-scale pretraining, inducing embedding space flattening. This challenges the conventional “fine-tuning always improves performance” paradigm in transfer learning—particularly on saturated benchmarks—and provides the first systematic evidence that excessive fine-tuning can impair retrieval ranking capability. Our findings suggest that architectural innovations or more robust adaptation mechanisms—not merely improved fine-tuning protocols—are critical to overcoming current limitations in dense retrieval.
This study challenges the implicit assumption that geometric proximity (e.g., cosine similarity) in sentence embedding spaces reflects semantic or functional similarity, asking whether such geometric properties can predict relative performance on downstream language tasks. Method: Within a unified Transformer framework, we systematically compare three embedding strategies—mean-pooled token, [CLS] token, and randomly selected token embeddings—across multiple NLP tasks. We conduct rigorous distance–performance correlation analysis to assess how well cosine similarity predicts task accuracy. Contribution/Results: We find that cosine similarity captures only shallow, surface-level lexical commonalities and fails to reliably predict downstream performance. Crucially, task-relevant semantic similarity is encoded via dimensionally weighted combinations rather than isotropic geometric proximity; thus, embeddings with large geometric distances in high-dimensional space may still encode highly similar task-specific semantics. This work provides the first empirical evidence of a substantial decoupling between the geometric structure of sentence embeddings and their functional utility, establishing a new paradigm for embedding evaluation and design.
Modeling the cosine similarity distribution of sentence embeddings remains challenging due to its bounded support and skewed, multimodal nature. Method: This work proposes, for the first time, a translated-and-truncated Gamma Mixture Model (GMM) constrained to [−1, 1] to characterize this distribution. Using a fixed corpus, it computes similarity scores between document embeddings and reference query embeddings, then fits the empirical density via expectation-maximization (EM). Contribution/Results: The model achieves high-fidelity distributional fitting across diverse corpora. Integrating hierarchical topic clustering, we uncover systematic correlations between distributional properties—such as kurtosis and skewness—and semantic topic hierarchies. This provides a novel statistical framework for semantic similarity modeling, enhances interpretability of embedding spaces, and releases an open-source, reusable modeling toolkit to support reproducible research in embedding analysis.
Vision-language models like CLIP align image and text embeddings but lack semantic comparability and analogical reasoning capabilities in their embedding space, hindering vectorized reasoning about inter-image differences. Method: We propose a contrastive learning fine-tuning framework that explicitly aligns image embedding differences to LLM-generated textual difference descriptors (e.g., “thinner”, “brighter”), enabling native pairwise image difference reasoning within CLIP’s embedding space. We further introduce a novel “comparative prompting” inference paradigm to restructure the embedding geometry. Contribution/Results: After fine-tuning on synthetic difference data, our method improves average accuracy by 3.2% across attribute ranking, zero-shot classification, and retrieval tasks. It significantly enhances linear analogy preservation and directional consistency in the embedding space—providing stronger geometric foundations for downstream applications such as text-to-image generation.
This work addresses the inefficiency of sentence embedding extraction from pretrained Transformers (e.g., BERT). We systematically investigate and enhance three key strategies: token aggregation, representation post-processing, and external-knowledge-guided fine-tuning. Specifically, we propose novel representation shaping techniques—including weighted aggregation of multi-layer hidden states, normalized contrastive fine-tuning, and Wikidata-augmented supervision—achieving substantial improvements in semantic expressiveness of static or randomly initialized embeddings, without introducing additional parameters or inference overhead. Our approach outperforms strong baselines across 8 semantic textual similarity, 6 short-text clustering, and 12 classification tasks. Notably, optimized random embeddings achieve over 120% improvement on STS-B, approaching native BERT performance. Empirical results validate the effectiveness and cross-model generalizability of lightweight representation shaping for universal sentence embedding learning.
This study investigates the capacity of sentence encoders to represent semantic concepts, revealing that existing models struggle to effectively learn relational and intensional concepts due to mismatches between architecture and supervision signals. Adopting a compositional representation perspective and leveraging a corpus of 3.3 million synonym-definition pairs, the work proposes four guiding principles: fine-tuning with recalibration outperforms expanding the latent space; semantic signals concentrate in the final Transformer layers; hard negative examples enhance discriminability without affecting ranking performance; and supervision efficacy depends on the compositional type of the concept. Through layer-wise pooling ablations, hard negative sampling, and training on large-scale lexical data, the authors construct a new evaluation benchmark—incorporating DBpedia and modifier-annotated noun phrases—and release two novel datasets, offering both theoretical insights and practical resources for research on conceptual representation.
Lightweight Transformer encoders for on-device multi-task NLP deployment face optimization conflicts and efficiency bottlenecks due to competing task gradients and stringent resource constraints. Method: We propose a task-guided LoRA-based multi-task pre-finetuning framework built upon a shared lightweight BERT backbone. It employs modular, task-specific LoRA adapters for named entity recognition (NER) and text classification, enabling parameter-efficient sharing while preserving task-specific optimization. A task-aware low-rank update mechanism mitigates gradient interference across tasks. Results: Evaluated on 21 downstream tasks, our method achieves +0.8% average F1 gain for NER and +8.8% accuracy improvement for text classification—matching the performance of standalone single-task fine-tuned models—while satisfying on-device deployment constraints in model size and computational cost. This work establishes a scalable, lightweight paradigm for edge-deployable multi-task NLP that balances generality, accuracy, and efficiency.
This work challenges the common practice in contrastive learning of using cosine similarity, which implicitly assumes that embedding norms are noise and discards their potential semantic information. Through a systematic 2×2 ablation study that independently controls input and output normalization in both text and vision encoders, the authors investigate the functional role of embedding norms. They propose a task symmetry principle: preserving norm information significantly improves performance in asymmetric tasks such as text retrieval, but harms performance in symmetric tasks. Furthermore, they reveal an asymmetric functional distinction between input and output norms. By combining controlled normalization, ablation experiments, and Cohen’s d effect size analysis, the study demonstrates that merely removing redundant unit hypersphere constraints at inference yields zero-cost performance gains on dense text retrieval benchmarks.
Embedding models trained independently on similar data capture stable semantic meanings but yield inconsistent representation spaces, hindering interoperability across models. This work addresses compatibility challenges in multimodal search and model upgrades via orthogonal transformation-based embedding alignment. Theoretically, we derive the first tight Procrustes alignment error bound, proving the existence of a near-isometric orthogonal transformation that approximately preserves pairwise inner products—establishing rigorous theoretical foundations for alignment. Methodologically, we employ efficient Procrustes analysis as a post-hoc alignment procedure, preserving the intrinsic geometric structure of each embedding space while enabling cross-model alignment. Experiments demonstrate substantial improvements in model retraining compatibility, text retrieval fusion accuracy, and cross-modal search performance; our method achieves state-of-the-art results in hybrid multimodal search.
This study systematically evaluates how text embedding models and retrieval pipeline configurations impact AI search performance, using high-quality evaluation data derived from U.S. city council meeting transcripts. Methodologically, it compares sentence-transformers models (All-MPNet, BGE, GTE) with a generative embedding model (Qwen3-Embedding-8B), incorporating fine-grained text chunking (512 characters), high-dimensional embeddings (4096 dimensions), Milvus indexing (HNSW/IVF), and neural re-ranking. It further introduces a local LLM-driven synthetic data generation framework and a CI/CD-automated evaluation pipeline. Key contributions include: (1) empirical validation that high-dimensional generative embeddings substantially improve long-tail query recall (Top-3 accuracy = 0.571); (2) identification of synergistic gains from fine-grained chunking and neural re-ranking; and (3) proposal of a reproducible, end-to-end optimized evaluation paradigm for AI search systems.