sbert finetuning

Fine-tune Sentence‑BERT / Sentence‑Transformers encoders to produce dense sentence embeddings whose vector geometry (cosine distance or dot product) encodes semantic similarity, using objectives such as cosine‑similarity regression, contrastive, or triplet losses. Design the finetuning pipeline — datasets and sampling, loss formulation, training and evaluation (e.g., retrieval recall@k, clustering, ranking) — and deliver models and embeddings for embedding‑based retrieval, clustering, or similarity scoring.

sbertfinetuning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.65
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

When Fine-Tuning Fails: Lessons from MS MARCO Passage Ranking

Jun 23, 2025
MP
Manu Pande
🏛️ IIIT Allahabad

This study identifies an anomalous performance degradation—where fine-tuning strong pretrained Transformer models (e.g., BERT, DeBERTa) on the MS MARCO passage ranking task reduces MRR@10 below the base model’s score (0.3026)—across all five fine-tuning strategies, including full-parameter tuning and parameter-efficient methods like LoRA. Using UMAP-based embedding visualization, training dynamics analysis, and efficiency evaluation, we demonstrate that fine-tuning disrupts the optimal semantic embedding structure acquired during large-scale pretraining, inducing embedding space flattening. This challenges the conventional “fine-tuning always improves performance” paradigm in transfer learning—particularly on saturated benchmarks—and provides the first systematic evidence that excessive fine-tuning can impair retrieval ranking capability. Our findings suggest that architectural innovations or more robust adaptation mechanisms—not merely improved fine-tuning protocols—are critical to overcoming current limitations in dense retrieval.

Challenges transfer learning effectiveness on saturated benchmarksFine-tuning degrades MS MARCO passage ranking performanceFine-tuning disrupts optimal pre-trained embedding space structure

Testing the assumptions about the geometry of sentence embedding spaces: the cosine measure need not apply

Sep 01, 2025
VN
Vivi Nastase
🏛️ Idiap Research Institute | University of Geneva

This study challenges the implicit assumption that geometric proximity (e.g., cosine similarity) in sentence embedding spaces reflects semantic or functional similarity, asking whether such geometric properties can predict relative performance on downstream language tasks. Method: Within a unified Transformer framework, we systematically compare three embedding strategies—mean-pooled token, [CLS] token, and randomly selected token embeddings—across multiple NLP tasks. We conduct rigorous distance–performance correlation analysis to assess how well cosine similarity predicts task accuracy. Contribution/Results: We find that cosine similarity captures only shallow, surface-level lexical commonalities and fails to reliably predict downstream performance. Crucially, task-relevant semantic similarity is encoded via dimensionally weighted combinations rather than isotropic geometric proximity; thus, embeddings with large geometric distances in high-dimensional space may still encode highly similar task-specific semantics. This work provides the first empirical evidence of a substantial decoupling between the geometric structure of sentence embeddings and their functional utility, establishing a new paradigm for embedding evaluation and design.

Evaluating cosine similarity's effectiveness for sentence embeddingsInvestigating if embedding distance predicts task performanceTesting geometry assumptions of sentence embedding spaces

Gamma Mixture Modeling for Cosine Similarity in Small Language Models

Oct 06, 2025
KP
Kevin Player
🏛️ Carnegie Mellon University

Modeling the cosine similarity distribution of sentence embeddings remains challenging due to its bounded support and skewed, multimodal nature. Method: This work proposes, for the first time, a translated-and-truncated Gamma Mixture Model (GMM) constrained to [−1, 1] to characterize this distribution. Using a fixed corpus, it computes similarity scores between document embeddings and reference query embeddings, then fits the empirical density via expectation-maximization (EM). Contribution/Results: The model achieves high-fidelity distributional fitting across diverse corpora. Integrating hierarchical topic clustering, we uncover systematic correlations between distributional properties—such as kurtosis and skewness—and semantic topic hierarchies. This provides a novel statistical framework for semantic similarity modeling, enhances interpretability of embedding spaces, and releases an open-source, reusable modeling toolkit to support reproducible research in embedding analysis.

Analyzing sentence transformer embeddings in small language modelsDeveloping expectation-maximization algorithm for shifted gamma mixturesModeling cosine similarity distributions using gamma mixtures

Finetuning CLIP to Reason about Pairwise Differences

Sep 15, 2024
DS
Dylan Sam
🏛️ Carnegie Mellon University | Bosch Center for AI

Vision-language models like CLIP align image and text embeddings but lack semantic comparability and analogical reasoning capabilities in their embedding space, hindering vectorized reasoning about inter-image differences. Method: We propose a contrastive learning fine-tuning framework that explicitly aligns image embedding differences to LLM-generated textual difference descriptors (e.g., “thinner”, “brighter”), enabling native pairwise image difference reasoning within CLIP’s embedding space. We further introduce a novel “comparative prompting” inference paradigm to restructure the embedding geometry. Contribution/Results: After fine-tuning on synthetic difference data, our method improves average accuracy by 3.2% across attribute ranking, zero-shot classification, and retrieval tasks. It significantly enhances linear analogy preservation and directional consistency in the embedding space—providing stronger geometric foundations for downstream applications such as text-to-image generation.

Enhance CLIP's ability to reason about image differencesEstablish geometric properties in CLIP's embedding spaceImprove zero-shot classification via comparative prompting

Extracting Sentence Embeddings from Pretrained Transformer Models

Aug 15, 2024
LS
Lukas Stankevicius
🏛️ Kaunas University of Technology

This work addresses the inefficiency of sentence embedding extraction from pretrained Transformers (e.g., BERT). We systematically investigate and enhance three key strategies: token aggregation, representation post-processing, and external-knowledge-guided fine-tuning. Specifically, we propose novel representation shaping techniques—including weighted aggregation of multi-layer hidden states, normalized contrastive fine-tuning, and Wikidata-augmented supervision—achieving substantial improvements in semantic expressiveness of static or randomly initialized embeddings, without introducing additional parameters or inference overhead. Our approach outperforms strong baselines across 8 semantic textual similarity, 6 short-text clustering, and 12 classification tasks. Notably, optimized random embeddings achieve over 120% improvement on STS-B, approaching native BERT performance. Empirical results validate the effectiveness and cross-model generalizability of lightweight representation shaping for universal sentence embedding learning.

Evaluates methods for extracting sentence embeddings from transformer models.Improves performance on Semantic Textual Similarity and clustering tasks.Tests token aggregation and post-processing techniques on BERT models.

Latest Papers

What's happening recently
View more

This study investigates the capacity of sentence encoders to represent semantic concepts, revealing that existing models struggle to effectively learn relational and intensional concepts due to mismatches between architecture and supervision signals. Adopting a compositional representation perspective and leveraging a corpus of 3.3 million synonym-definition pairs, the work proposes four guiding principles: fine-tuning with recalibration outperforms expanding the latent space; semantic signals concentrate in the final Transformer layers; hard negative examples enhance discriminability without affecting ranking performance; and supervision efficacy depends on the compositional type of the concept. Through layer-wise pooling ablations, hard negative sampling, and training on large-scale lexical data, the authors construct a new evaluation benchmark—incorporating DBpedia and modifier-annotated noun phrases—and release two novel datasets, offering both theoretical insights and practical resources for research on conceptual representation.

concept representationrepresentational compositionalitysemantic operators

Multi-Task Pre-Finetuning of Lightweight Transformer Encoders for Text Classification and NER

Oct 08, 2025
JZ
Junyi Zhu
🏛️ Samsung R&D Institute UK | Samsung Electronics Korea

Lightweight Transformer encoders for on-device multi-task NLP deployment face optimization conflicts and efficiency bottlenecks due to competing task gradients and stringent resource constraints. Method: We propose a task-guided LoRA-based multi-task pre-finetuning framework built upon a shared lightweight BERT backbone. It employs modular, task-specific LoRA adapters for named entity recognition (NER) and text classification, enabling parameter-efficient sharing while preserving task-specific optimization. A task-aware low-rank update mechanism mitigates gradient interference across tasks. Results: Evaluated on 21 downstream tasks, our method achieves +0.8% average F1 gain for NER and +8.8% accuracy improvement for text classification—matching the performance of standalone single-task fine-tuned models—while satisfying on-device deployment constraints in model size and computational cost. This work establishes a scalable, lightweight paradigm for edge-deployable multi-task NLP that balances generality, accuracy, and efficiency.

Developing efficient NLP models for mobile deployment with limited resourcesEnhancing lightweight encoders for both text classification and NER tasksResolving conflicting optimization signals in multi-task pre-finetuning

This work challenges the common practice in contrastive learning of using cosine similarity, which implicitly assumes that embedding norms are noise and discards their potential semantic information. Through a systematic 2×2 ablation study that independently controls input and output normalization in both text and vision encoders, the authors investigate the functional role of embedding norms. They propose a task symmetry principle: preserving norm information significantly improves performance in asymmetric tasks such as text retrieval, but harms performance in symmetric tasks. Furthermore, they reveal an asymmetric functional distinction between input and output norms. By combining controlled normalization, ablation experiments, and Cohen’s d effect size analysis, the study demonstrates that merely removing redundant unit hypersphere constraints at inference yields zero-cost performance gains on dense text retrieval benchmarks.

contrastive learningcosine similaritydot product

When Embedding Models Meet: Procrustes Bounds and Applications

Oct 15, 2025
LM
Lucas Maystre
🏛️ UiPath | Spotify

Embedding models trained independently on similar data capture stable semantic meanings but yield inconsistent representation spaces, hindering interoperability across models. This work addresses compatibility challenges in multimodal search and model upgrades via orthogonal transformation-based embedding alignment. Theoretically, we derive the first tight Procrustes alignment error bound, proving the existence of a near-isometric orthogonal transformation that approximately preserves pairwise inner products—establishing rigorous theoretical foundations for alignment. Methodologically, we employ efficient Procrustes analysis as a post-hoc alignment procedure, preserving the intrinsic geometric structure of each embedding space while enabling cross-model alignment. Experiments demonstrate substantial improvements in model retraining compatibility, text retrieval fusion accuracy, and cross-modal search performance; our method achieves state-of-the-art results in hybrid multimodal search.

Aligning embeddings from separately trained modelsEnabling interoperability through orthogonal transformationsImproving multimodal search and model compatibility

Evaluating Embedding Models and Pipeline Optimization for AI Search Quality

Nov 27, 2025
PZ
Philip Zhong
🏛️ Cisco Systems, Inc.

This study systematically evaluates how text embedding models and retrieval pipeline configurations impact AI search performance, using high-quality evaluation data derived from U.S. city council meeting transcripts. Methodologically, it compares sentence-transformers models (All-MPNet, BGE, GTE) with a generative embedding model (Qwen3-Embedding-8B), incorporating fine-grained text chunking (512 characters), high-dimensional embeddings (4096 dimensions), Milvus indexing (HNSW/IVF), and neural re-ranking. It further introduces a local LLM-driven synthetic data generation framework and a CI/CD-automated evaluation pipeline. Key contributions include: (1) empirical validation that high-dimensional generative embeddings substantially improve long-tail query recall (Top-3 accuracy = 0.571); (2) identification of synergistic gains from fine-grained chunking and neural re-ranking; and (3) proposal of a reproducible, end-to-end optimized evaluation paradigm for AI search systems.

Comparing chunking strategies and indexing methodsEvaluating embedding models for AI search qualityOptimizing pipeline configurations for retrieval accuracy

Hot Scholars

KH

Kuo-Hui Yeh

National Yang Ming Chiao Tung University
SecurityPrivacy
AT

Abhiroop Talasila

International Institute of Information Technology, Hyderabad
AI applications for SGDsHealthcare & AI
SR

Shivam Ratnakar

Graduate student at USC Viterbi
Machine LearningNLPGenerative SearchRecommendation Engines
HS

Hajar Sakai

Ph.D. in Industrial and Systems Engineering
Large Language ModelsText ClassificationTime Series Forecasting
RJ

Raviraj Joshi

Indian Institute of Technology Madras
computer sciencemachine learningnatural language processing