train sentence transformers

Designs and trains transformer-based models that produce dense sentence- or document-level embeddings by applying contrastive and document-level fine-tuning techniques; builds and evaluates these sentence-transformer models to improve downstream prediction accuracy, robustness to input perturbations, and to capture linguistic signals (e.g., hedging).

trainsentencetransformers

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

When Text Embedding Meets Large Language Model: A Comprehensive Survey

Dec 12, 2024
ZN
Zhijie Nie
🏛️ Beihang University

This work addresses the joint optimization of large language models (LLMs) and text embedding techniques to enhance efficiency and robustness in semantic matching, clustering, and information retrieval. We propose the first unified taxonomy centered on the *interaction patterns* between LLMs and embeddings—categorizing approaches into three paradigms: LLM-augmented embeddings, LLM-as-embedder, and LLM-understanding-embeddings—thereby transcending conventional task-centric taxonomies. By integrating supervised/unsupervised embedding learning, instruction tuning, prompt engineering, representation space analysis, and interpretability methods, we construct a structured knowledge graph encompassing over 100 studies. Our framework precisely delineates capability boundaries and application scopes for each paradigm, identifies persistent limitations inherited from pre-trained language models (PLMs) and novel challenges introduced by LLMs, and provides a theoretically grounded, empirically informed roadmap for future advancement.

Analyzing and interpreting text embeddings with LLMsCombining LLMs and text embeddings for NLP advancementsEnhancing text embedding methods using large language models

Must-Read Papers

Most classic and influential ideas
View more

Eliciting Fine-Tuned Transformer Capabilities via Inference-Time Techniques

Jun 09, 2025
AS
Asankhaya Sharma
🏛️ Patched Codes, Inc.

This work addresses the challenge of enabling base large language models to approximate supervised fine-tuning (SFT) performance at inference time—without any parameter updates or SFT. Method: We propose a parameter-free approach grounded in in-context learning (ICL), retrieval-augmented generation, and probabilistic generalization theory. Contribution/Results: We provide the first theoretical proof that Transformer-based base models can approximate arbitrary SFT policies using only a finite context window and a bounded number of in-context examples. Leveraging the Turing completeness of Transformers, we derive sample complexity upper bounds for both text generation and linear classification tasks. Our analysis establishes the first solution that simultaneously offers rigorous theoretical guarantees and practical deployment feasibility—enabling low-overhead, fine-tuning-free model adaptation.

Approximating fine-tuned transformer capabilities without parameter updatesReducing computational costs of supervised fine-tuning via inference-time techniquesTheoretical bounds for dataset sizes to achieve fine-tuned behavior

Extracting Sentence Embeddings from Pretrained Transformer Models

Aug 15, 2024
LS
Lukas Stankevicius
🏛️ Kaunas University of Technology

This work addresses the inefficiency of sentence embedding extraction from pretrained Transformers (e.g., BERT). We systematically investigate and enhance three key strategies: token aggregation, representation post-processing, and external-knowledge-guided fine-tuning. Specifically, we propose novel representation shaping techniques—including weighted aggregation of multi-layer hidden states, normalized contrastive fine-tuning, and Wikidata-augmented supervision—achieving substantial improvements in semantic expressiveness of static or randomly initialized embeddings, without introducing additional parameters or inference overhead. Our approach outperforms strong baselines across 8 semantic textual similarity, 6 short-text clustering, and 12 classification tasks. Notably, optimized random embeddings achieve over 120% improvement on STS-B, approaching native BERT performance. Empirical results validate the effectiveness and cross-model generalizability of lightweight representation shaping for universal sentence embedding learning.

Evaluates methods for extracting sentence embeddings from transformer models.Improves performance on Semantic Textual Similarity and clustering tasks.Tests token aggregation and post-processing techniques on BERT models.

Physics of Language Models: Part 1, Learning Hierarchical Language Structures

May 23, 2023
ZA
Zeyuan Allen-Zhu
🏛️ Meta | FAIR Labs | Mohamed bin Zayed University of AI

This work investigates how Transformer language models capture deep recursive structures defined by context-free grammars (CFGs). To address their limited reasoning capability on long-range, ambiguous nested sentences, we construct controllable synthetic CFG families that generate challenging long sequences, and conduct systematic analysis via hidden-state interpretability, attention visualization, and analogy to dynamic programming. Our key findings are: (1) generative models (e.g., GPT) precisely encode syntactic tree depth in hidden states, and their attention patterns explicitly emulate dynamic programming steps; (2) standard positional encodings suffer from representational degradation under deep nesting; (3) generative architectures substantially outperform encoder-only models (e.g., BERT, deBERTa). Building on these insights, we propose a structured error pretraining strategy that significantly enhances robust CFG modeling—achieving strong generalization to sequences exceeding hundreds of tokens.

Comparing model performance on deep structure reasoning across architecturesInvestigating hidden states and attention patterns in models processing CFGsUnderstanding how language models learn hierarchical structures defined by CFGs

This study systematically investigates whether and how Transformer-based language models acquire syntactic knowledge. Through a large-scale, systematic literature review synthesizing findings from 337 studies and over 3,000 data points, the work presents the first integrated quantitative assessment of syntactic capabilities across multiple languages and model architectures by combining behavioral experiments, representation probing, and mechanistic interpretability methods. The analysis reveals that Transformers possess substantial syntactic knowledge, yet exhibit limitations in phenomena at the syntax–semantics interface and in low-resource languages. It also highlights a pronounced research bias toward English and BERT-family models, with insufficient coverage of linguistic and architectural diversity. This work provides comprehensive empirical evidence and new directions for understanding the mechanisms and boundaries of syntactic generalization in neural language models.

cross-lingualinterpretabilitysyntactic knowledge

Learning Syntax Without Planting Trees: Understanding When and Why Transformers Generalize Hierarchically

Apr 25, 2024
KA
Kabir Ahuja
🏛️ University of Washington | Carnegie Mellon University | Microsoft Research | Allen Institute for AI

This study investigates the mechanisms enabling Transformers to achieve hierarchical generalization without explicit syntactic supervision, and identifies the sources of their inductive biases. Method: We employ synthetic syntactic datasets, multi-objective contrastive training (language modeling, prefix-LM, and seq2seq), structured pruning, and Bayesian model selection analysis. Contribution/Results: (1) Standard language modeling (LM) objective alone suffices to robustly induce hierarchical generalization—serving as the critical training signal. (2) Transformers internally host parallel, co-existing hierarchical and linear subnetworks. (3) The propensity for hierarchical generalization aligns closely with the Bayesian principle of “simplest grammatical explanation”: Transformers consistently exhibit hierarchical behavior precisely when hierarchical grammars minimize description length. This work establishes, for the first time, a causal chain from the LM objective → internal dual-path architecture → Bayesian-optimal syntactic explanation.

Analyzes subnetworks in transformers for hierarchical and linear generalization.Explores training objectives enabling hierarchical generalization in transformers.Investigates inductive bias in transformers for hierarchical generalization.

Latest Papers

What's happening recently
View more

To address the lack of high-quality, controllable sentence-level embeddings from large language models (LLMs) for non-generative tasks—such as clustering, classification, and retrieval—this paper proposes a unified framework integrating prompt engineering, contrastive fine-tuning, and semantic-aware aggregation. Specifically, it introduces task-oriented prompt templates, synthesizes positive pairs to drive contrastive learning, and employs attention-based token-level vector weighting for sentence embedding aggregation—thereby preserving salient semantics while suppressing noise. The method requires only lightweight fine-tuning of decoder-only LLMs and is validated via attention analysis to confirm enhanced semantic focus. Evaluated on the MTEB English clustering benchmark, it achieves state-of-the-art performance, significantly outperforming mainstream embedding models. Results demonstrate the framework’s effectiveness, robustness, and practical utility for sentence embedding generation in non-generative downstream applications.

Adapting LLMs for non-generative tasks via efficient methodsEnhancing semantic compression in embeddings via contrastive fine-tuningImproving text embeddings from token-level LLM representations

This study addresses the challenge of aggregating multiple texts from individuals for precise mental health assessment by systematically comparing base Transformers with document-level fine-tuned Transformers on longitudinal psychological data. It provides the first empirical validation of the advantages of document-level fine-tuned representations for this task. Through contrastive learning–based fine-tuning, layer-wise representation analysis, and robustness evaluations under various linguistic perturbations—including word deletion, synonym substitution, spelling errors, and back-translation—the fine-tuned models demonstrate significantly superior performance over baseline architectures, achieving a 13.4% improvement in Pearson correlation coefficient (p = 0.015). Moreover, these models exhibit enhanced capability in capturing linguistic uncertainty and greater robustness to textual variations.

document-level representationmental health assessmentnatural language processing

This study addresses the computational bottlenecks of byte-level language models caused by excessively long sequences and the absence of explicit textual abstraction. To overcome these limitations, this work proposes a tokenizer-free architecture that leverages a Token Hyper-position training strategy alongside hash embeddings, enabling standard Transformers to model efficiently at the byte level. The findings demonstrate that additional computation can effectively substitute fixed tokenizers, allowing models to spontaneously construct local context representation mechanisms. Furthermore, the proposed approach surpasses subword-based models in performance at large scales. Notably, by exploiting the non-uniform uncertainty inherent in generated outputs, the method facilitates speculative decoding, achieving a 3.4-fold improvement in acceptance rates.

Byte Language ModelsEmergent AbstractionsTokenizer-free

This work addresses the rapid yet often unstructured evolution of Transformer-based language models by proposing a practical, four-dimensional evaluation framework to distinguish substantive advances from incremental improvements and to guide cross-domain deployment. The framework holistically assesses model architecture, alignment methodologies, energy-efficiency trade-offs, and domain adaptability, encompassing key techniques such as encoder/decoder variants, long-context modeling, mixture-of-experts (MoE), retrieval augmentation, instruction tuning, and preference optimization. Through empirical analysis across vertical domains—including healthcare, finance, and law—the study quantifies the trade-offs between model scale and computational cost, redefines what constitutes an “advanced” model in real-world settings, identifies critical research gaps, and offers actionable guidelines for model selection and deployment.

architectural comparisondomain applicationsmodel evaluation

Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions

Sep 04, 2025
FA
Faruk Alpay
🏛️ Lightcap | Turkish Aeronautical Association

Transformer language models lack fine-grained controllability for tasks such as constrained text generation, behavioral alignment, and robust intervention. Method: This paper introduces the first unified three-tier intervention framework—prompt guidance, activation intervention, and weight editing—integrating prompt engineering, parameter-efficient fine-tuning, model editing, and reinforcement learning to enable targeted behavioral modulation without degrading core capabilities. Contribution/Results: We theoretically establish that small-magnitude weight updates suffice for low-side-effect behavioral customization and formally characterize a safe intervention boundary. Empirical evaluation demonstrates >90% success rates in sentiment control and factual correction tasks, revealing fundamental trade-offs between generality and specificity. The framework provides a verifiable, interpretable, and principled paradigm for controllable AI.

Achieving fine-grained control in transformer-based language modelsAnalyzing robustness and safety implications of model interventionsFormalizing controllable text generation as optimization problem

Hot Scholars

GZ

Guorui Zhou

Unknown affiliation
Recommender System,Advertising,Artificial Intelligence,Machine Learning,NLP
BS

Biplav Srivastava

University of South Carolina
Artificial Intelligence (Automated PlanningTrustLearning)Smarter Cities (Water)
NG

Nuno Gracias

Computer Vision and Robotics Institute, University of Girona
Computer VisionRoboticsImage processingUnderwater Image Processing
XZ

Xueyao Zhang

The Chinese University of Hong Kong, Shenzhen
Deep LearningSpeechSinging VoiceMusic