two-stage decoder training

Designs and implements training procedures and optimization schedules that first train a shared model component and then specialize one or more decoders (or transition from a shared decoder to multiple specialized decoders). Builds the transition mechanics, parameter transfer or fine-tuning steps, and evaluation/diagnostics to close the performance gap with fully specialized models and to stabilize learning across different latency or configuration regimes.

two-stagedecodertraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.03
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work investigates whether pretraining of large language models can be decomposed into independent smaller-scale tasks and later reassembled into a fully functional model. To this end, the authors propose Mixture of Training (MoT), a method that partitions a target Transformer into contiguous layer blocks, trains these blocks in parallel within a frozen pretrained aligner scaffold, and subsequently reassembles them followed by brief end-to-end fine-tuning. Evaluated on the Gemma architecture using the C4 dataset, MoT achieves perplexity comparable to monolithic training on a 1.3B-parameter model despite processing more total tokens, and demonstrates potential computational efficiency when reusing the aligner. This study provides the first empirical validation that deep architectural slices can be independently trained and effectively recombined into a coherent large language model.

language model pretrainingmodel decompositionmodular training

Learning rate scheduling in large language model training lacks rigorous theoretical foundations, leading to heuristic designs and suboptimal convergence. Method: This paper establishes, for the first time, a quantitative alignment between practical schedulers (e.g., linear decay) and tight non-smooth convex optimization lower bounds—eliminating spurious logarithmic factors in prior analyses and enabling principled cross-scheduler optimal learning rate transfer. We integrate convex optimization theory, scheduler modeling, and empirical validation, conducting systematic evaluations on 124M- and 210M-parameter Llama models. Results: Theory-guided scheduler design yields faster convergence and improved stability, empirically validating optimization theory’s practical relevance for large-model training. Core contribution: bridging the gap between theoretical performance bounds and engineering schedulers by providing a transferable, interpretable, and theoretically grounded framework for learning rate tuning.

Large Model TrainingLearning Rate AdjustmentTraining Efficiency

In large-scale pretraining, learning rate scheduling critically influences both training efficiency and model performance. This work proposes two paradigms—Fitting and Transfer. The Fitting paradigm establishes, for the first time, a scaling law for learning rate search factors, reducing hyperparameter tuning complexity from O(n³) to O(n·C_D·C_η). The Transfer paradigm extends μTransfer to Mixture-of-Experts (MoE) architectures and generalizes it across multiple hyperparameter dimensions, including depth, weight decay, and token length. Empirical results demonstrate that while μTransfer exhibits limited scalability in large-scale settings, the Fitting paradigm—grounded in the derived scaling law—offers superior scalability and practicality, providing a systematic guideline for hyperparameter tuning in industrial-scale pretraining.

hyperparameter optimizationlarge-scale pre-traininglearning rate

MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parameters

Feb 04, 2024
AS
Arsalan Sharifnassab
🏛️ University of Alberta | Leiden University

To address the inefficiency and poor generalizability of manual hyperparameter tuning—particularly for learning rates—this paper proposes a dynamic online meta-optimization framework that formulates learning rate adaptation as a discounted cumulative regret minimization problem over time. The method employs a gradient-based meta-update mechanism, enabling plug-and-play integration with any first-order optimizer (e.g., SGD, Adam) to achieve decoupled, real-time, adaptive step-size optimization. Key contributions include: (i) the first formalization of meta-optimization as discounted regret minimization; and (ii) a low-complexity variant that preserves theoretical rigor while ensuring computational efficiency and strong generalization. Experiments across diverse tasks demonstrate faster convergence, enhanced robustness to initialization and task heterogeneity, competitive performance against hand-tuned optimal schedulers, and significantly lower computational overhead compared to conventional hyperparameter search methods.

Dynamically adjusting step sizes during model optimizationOptimizing meta-parameters for efficient machine learning trainingReducing regret by considering long-term impact of learning rates

STEP: Staged Parameter-Efficient Pre-training for Large Language Models

Apr 05, 2025
KY
Kazuki Yano
🏛️ Tohoku University | Langsmith Inc. | RIKEN | NII LLMC

To address the prominent GPU memory bottleneck in large language model (LLM) pretraining, this paper proposes the Staged Parameter-Efficient Training (SPET) framework. SPET is the first to deeply integrate parameter-efficient fine-tuning techniques—such as LoRA—into the *entire* pretraining pipeline, synergistically combining gradient checkpointing with staged architectural expansion to enable dynamic model growth and on-demand memory optimization. Implemented in PyTorch, SPET introduces a memory-aware training scheduler that reduces peak GPU memory consumption by up to 53.9% versus full-parameter baselines, while preserving pretraining performance. Downstream task performance after instruction tuning remains unchanged. The core contribution lies in bridging the paradigmatic divide between standard pretraining and parameter-efficient adaptation, establishing a scalable, memory-efficient, and unified pretraining paradigm.

Integrates efficient tuning with model growthMaintains performance with less memoryReduces memory use in LLM pre-training

Latest Papers

What's happening recently
View more

This study addresses the observation that the generalization capability of language models during pretraining does not improve monotonically, but instead oscillates frequently between rote memorization and intelligent reasoning. To investigate this, we construct an evaluation suite to identify and define the "mode jumping" phenomenon, modeling it as a circuit competition problem under capacity constraints. We propose a theoretical framework for capacity allocation, wherein data windows govern circuit competition, and integrate intermediate checkpoint selection with pretraining data selection strategies to monitor and control generalization dynamics. Our findings challenge the conventional assumption of stable model maturation by demonstrating that intermediate checkpoints can exhibit superior reasoning and alignment capabilities compared to the final model. Furthermore, we show that strategic data selection effectively stabilizes the generalization process throughout pretraining.

Capacity AllocationGeneralization DynamicsLanguage Model Pre-training

This study addresses the challenges of indeterminate skill bottleneck resolution order and inefficient data mixing in large language model training. We propose a staged training framework integrating small proxy model exploration with a LogFloor closed-loop controller. This approach transforms bottleneck resolution trajectories into transferable curriculum learning structures, employing a "small-model reconnaissance and path transfer" mechanism to guide large models through sequential bottleneck breakthroughs. Experiments on Qwen2.5 demonstrate that this strategy reduces training tokens by an average of 56.2% and achieves approximately 39% computational savings through cross-scale transfer. These results indicate significant improvements in data efficiency for skill acquisition in large language models, offering a scalable solution to optimize training dynamics and resource utilization.

Data ControlProxy ModelSkill Bottleneck

This study addresses the limitation that transferring capabilities from expert models to general-purpose language models typically relies on training or alignment. To overcome this, we propose a training-free heterogeneous model merging method that eliminates the need for gradient updates and semantic alignment. By leveraging parameter projection and interpolation, our approach directly facilitates cross-role knowledge transfer at the parameter level. Furthermore, we design two core merging strategies: Intersection-Merge and Activate-Prune-Merge. Experimental results demonstrate that the proposed method significantly enhances the performance of general-purpose models across embedding, reranking, and code generation tasks. This work establishes a novel paradigm for the efficient integration of heterogeneous models.

Heterogeneous Model MergingKnowledge TransferSpecialist-to-General Transfer

Hot Scholars

TG

Tianpei Gu

Research Scientist, ByteDance/TikTok
Computer VisionGenerative Model
BW

Bihan Wen

Associate Professor, Nanyang Technological University
Machine LearningImage ProcessingComputational ImagingComputer Vision
BT

Benedetta Tondi

Department of Information Engineering of the University of Siena
Image ForensicsSignal processingMultimedia securityInformation theory
YQ

Yulei Qin

Tencent YouTu Lab
Language ModelsComputer VisionMedical Image Analysis