fine-tune embeddings

Design and implement procedures, adapter layers, or transform pipelines that adapt pre-trained embedding representations to a target dataset or downstream predictor by updating embedding weights or learning lightweight modifications to improve downstream predictive accuracy. Build training, evaluation, and deployment practices that address limited-data regimes and that ensure deterministic, reproducible embedding extraction (fixed checkpoints, seed control, or deterministic transforms) for governance and downstream model reliability.

fine-tuneembeddings

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.15
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

AdaPtis: Reducing Pipeline Bubbles with Adaptive Pipeline Parallelism on Heterogeneous Models

Sep 28, 2025
JG
Jihu Guo
🏛️ Fudan University | Shanghai AI Laboratory | Hong Kong University of Science and Technology | SenseTime | Tsinghua University | Chinese University of Hong Kong | Sensetime Research

To address severe pipeline bubbles and limited throughput in heterogeneous large-model training, this paper proposes AdaPtis, an adaptive pipeline parallelism system. Methodologically, AdaPtis introduces (1) a generalizable pipeline performance model; (2) the first joint optimization of model partitioning, device placement, and micro-batch scheduling; and (3) a unified pipeline executor supporting diverse parallelism strategies. Experiments on representative heterogeneous hardware configurations demonstrate that AdaPtis achieves an average 1.42× speedup over Megatron-LM’s I-1F1B baseline, with peak improvements reaching 2.14×. These gains translate into significantly enhanced training efficiency and improved hardware resource utilization, without compromising model accuracy or training stability.

Co-optimizes model partition, placement and workload schedulingImproves training efficiency through adaptive pipeline parallelismReduces pipeline bubbles in heterogeneous LLM training

Existing theoretical frameworks struggle to explain why larger-scale pre-trained models substantially reduce sample complexity on downstream tasks. This work proposes a novel theoretical framework—termed “caulking”—inspired by parameter-efficient fine-tuning methods such as adapters, low-rank adaptation, and partial fine-tuning. It establishes, for the first time, a provable relationship between the scale of pre-trained models and the sample complexity of downstream tasks. By rigorously linking stronger pre-training capabilities to reduced data requirements in transfer learning, this study not only addresses a critical gap in current theoretical understanding but also provides a solid foundation for empirically observed scaling laws, demonstrating that enhanced pre-training capacity can significantly decrease the number of samples needed for effective downstream adaptation.

downstream taskspre-trained modelssample complexity

This study addresses the over-reliance of language agents on external frameworks for state tracking and control decision-making by proposing Harness Annealing Training, a method designed to internalize control capabilities. Introducing the novel concept of harness annealing, this approach integrates explicit control supervision with a progressive attenuation strategy. Through curriculum learning on teacher trajectories, it distills successful trajectories into the model’s intrinsic decision-making capacity while dynamically adjusting the allocation of control authority between human and machine. Experimental results demonstrate that 9B- and 35B-parameter models, when relying solely on tool use, achieve performance comparable to full framework deployment, thereby significantly reducing runtime control overhead.

autonomous decision-makingcontrol decisionsexternal control

Latest Papers

What's happening recently
View more

This work addresses the fragmented landscape of post-training adaptation techniques, which suffer from inconsistent terminology and a lack of unified comparative or governance frameworks. To resolve this, the paper introduces the first six-dimensional taxonomy—spanning mechanism, objective, data requirements, persistence, structural scope, and model type—that systematically integrates mainstream approaches such as fine-tuning, retrieval augmentation, prompt engineering, model editing, and machine unlearning. This framework clarifies conceptual boundaries and reveals evolutionary and compositional relationships among methods. Beyond standardizing terminology, it enables standardized technical documentation, model change tracking, and AI governance analysis. The study further identifies critical challenges, including evaluation rigor, reproducibility, continual adaptation, multimodal alignment, and governance-aware workflows.

AI governancefoundation modelsmodel modification

This work addresses the vulnerability of open-source large language models (LLMs) to malicious prompts—even on harmless inputs—caused by the degradation of safety alignment during downstream fine-tuning. To mitigate this, the authors propose SafeGene, a reusable safety adapter module that, for the first time, explicitly models safety alignment as a task-agnostic, standalone representation. SafeGene extracts safety vectors by contrasting aligned and degraded models, and integrates a data-aware layer selection strategy with few-shot, layer-wise coefficient calibration. Evaluated across multiple model families, downstream tasks, and safety benchmarks, SafeGene significantly reduces harmful response rates while preserving task performance, outperforming existing safety adaptation methods.

downstream fine-tuningharmful promptsmodel safety

This study addresses the absence of supervision signals caused by the unavailability of specialized testing frameworks during deployment. To overcome this limitation, we propose a recursive self-rewriting framework built upon the Qwen-3.8-27B model. By orchestrating the collaboration of a planner, a critic, and an executor, our approach reconstructs problem-solving experiences from multiple frameworks into generalized training trajectories. Furthermore, reliability is ensured through runbook extraction, leakage screening, and sandboxed execution techniques. Experimental results demonstrate that the proposed method improves the task resolution rate by 34.3% and significantly outperforms baselines on the Pass@3 metric across benchmarks such as Terminal-Bench. These findings confirm that our framework effectively enables the reuse and generalization of cross-framework experiences in practical deployment scenarios.

complex tasksdeployment gapharness generalization

This study addresses the challenges of missing source data and catastrophic forgetting during fine-tuning when transferring pretrained chip placement models across designs. We propose CARVE, a continual adaptation framework that enables secure knowledge reuse through frozen base policies, immutable expert modules, and task credential mechanisms. Furthermore, we introduce a novel local-validation-based expert selection strategy, establishing theoretical bounds for expected performance guarantees under fixed distributions and worst-case sample complexity with bounded loss. Experimental results demonstrate that our method reduces training time by 58.5% while preserving nearly lossless HPWL gains. Notably, on IBM circuits, CARVE achieves an average 5.76% HPWL improvement without requiring any receiver-side training.

Catastrophic ForgettingChip PlacementContinual Adaptation

This work proposes DataOrchestra, a novel framework that addresses the limitations of conventional pretraining data processing, which typically applies uniform strategies and fails to account for sample-level heterogeneity. DataOrchestra introduces the first automated, per-sample orchestration mechanism, featuring a hierarchical decision architecture: a high-level orchestrator dynamically selects among discarding, retaining, or combining multiple cleaning operations—such as programmatic editing and large language model–based rewriting—while low-level tool models execute the chosen operations accordingly. Evaluated across 11 benchmark tasks on models ranging from 0.5B to 7B parameters, this approach consistently outperforms any single static processing strategy, demonstrating particularly strong gains in mathematical continual pretraining while substantially reducing computational overhead.

adaptive data processingdata cleaninglarge language models

Hot Scholars

YK

Ye Kyaw Thu

LST Lab., NECTEC (Thai), NLP Research Lab, UTYCC (Myanmar), Language Understanding Lab., (Myanmar)
Natural Language ProcessingMachine TranslationSpeech ProcessingAI
AP

Aritran Piplai

The University of Texas at El Paso
Artificial intelligenceKnowledge extractioncyber security
ZX

Ziqi Xu

Lecturer, School of Computing Technologies, RMIT University
Causal AIFairness
XG

Xiaobo Guo

Dartmouth College
machine learningdeep learningnatural language processingsocia media
CF

Changjun Fan

Associate Professor, National University of Defense Technology
graph neural networkcombinatorial optimizationreinforcement learning