Score
Designs and implements training procedures and optimization schedules that first train a shared model component and then specialize one or more decoders (or transition from a shared decoder to multiple specialized decoders). Builds the transition mechanics, parameter transfer or fine-tuning steps, and evaluation/diagnostics to close the performance gap with fully specialized models and to stabilize learning across different latency or configuration regimes.
This work investigates whether pretraining of large language models can be decomposed into independent smaller-scale tasks and later reassembled into a fully functional model. To this end, the authors propose Mixture of Training (MoT), a method that partitions a target Transformer into contiguous layer blocks, trains these blocks in parallel within a frozen pretrained aligner scaffold, and subsequently reassembles them followed by brief end-to-end fine-tuning. Evaluated on the Gemma architecture using the C4 dataset, MoT achieves perplexity comparable to monolithic training on a 1.3B-parameter model despite processing more total tokens, and demonstrates potential computational efficiency when reusing the aligner. This study provides the first empirical validation that deep architectural slices can be independently trained and effectively recombined into a coherent large language model.
Learning rate scheduling in large language model training lacks rigorous theoretical foundations, leading to heuristic designs and suboptimal convergence. Method: This paper establishes, for the first time, a quantitative alignment between practical schedulers (e.g., linear decay) and tight non-smooth convex optimization lower bounds—eliminating spurious logarithmic factors in prior analyses and enabling principled cross-scheduler optimal learning rate transfer. We integrate convex optimization theory, scheduler modeling, and empirical validation, conducting systematic evaluations on 124M- and 210M-parameter Llama models. Results: Theory-guided scheduler design yields faster convergence and improved stability, empirically validating optimization theory’s practical relevance for large-model training. Core contribution: bridging the gap between theoretical performance bounds and engineering schedulers by providing a transferable, interpretable, and theoretically grounded framework for learning rate tuning.
In large-scale pretraining, learning rate scheduling critically influences both training efficiency and model performance. This work proposes two paradigms—Fitting and Transfer. The Fitting paradigm establishes, for the first time, a scaling law for learning rate search factors, reducing hyperparameter tuning complexity from O(n³) to O(n·C_D·C_η). The Transfer paradigm extends μTransfer to Mixture-of-Experts (MoE) architectures and generalizes it across multiple hyperparameter dimensions, including depth, weight decay, and token length. Empirical results demonstrate that while μTransfer exhibits limited scalability in large-scale settings, the Fitting paradigm—grounded in the derived scaling law—offers superior scalability and practicality, providing a systematic guideline for hyperparameter tuning in industrial-scale pretraining.
To address the inefficiency and poor generalizability of manual hyperparameter tuning—particularly for learning rates—this paper proposes a dynamic online meta-optimization framework that formulates learning rate adaptation as a discounted cumulative regret minimization problem over time. The method employs a gradient-based meta-update mechanism, enabling plug-and-play integration with any first-order optimizer (e.g., SGD, Adam) to achieve decoupled, real-time, adaptive step-size optimization. Key contributions include: (i) the first formalization of meta-optimization as discounted regret minimization; and (ii) a low-complexity variant that preserves theoretical rigor while ensuring computational efficiency and strong generalization. Experiments across diverse tasks demonstrate faster convergence, enhanced robustness to initialization and task heterogeneity, competitive performance against hand-tuned optimal schedulers, and significantly lower computational overhead compared to conventional hyperparameter search methods.
To address the prominent GPU memory bottleneck in large language model (LLM) pretraining, this paper proposes the Staged Parameter-Efficient Training (SPET) framework. SPET is the first to deeply integrate parameter-efficient fine-tuning techniques—such as LoRA—into the *entire* pretraining pipeline, synergistically combining gradient checkpointing with staged architectural expansion to enable dynamic model growth and on-demand memory optimization. Implemented in PyTorch, SPET introduces a memory-aware training scheduler that reduces peak GPU memory consumption by up to 53.9% versus full-parameter baselines, while preserving pretraining performance. Downstream task performance after instruction tuning remains unchanged. The core contribution lies in bridging the paradigmatic divide between standard pretraining and parameter-efficient adaptation, establishing a scalable, memory-efficient, and unified pretraining paradigm.
研究探讨了客户服务LLM的多任务处理策略,通过多任务微调、顺序更新或模型合并的方法,发现多任务全微调在所有测试模型尺寸中表现最佳。
This study addresses the observation that the generalization capability of language models during pretraining does not improve monotonically, but instead oscillates frequently between rote memorization and intelligent reasoning. To investigate this, we construct an evaluation suite to identify and define the "mode jumping" phenomenon, modeling it as a circuit competition problem under capacity constraints. We propose a theoretical framework for capacity allocation, wherein data windows govern circuit competition, and integrate intermediate checkpoint selection with pretraining data selection strategies to monitor and control generalization dynamics. Our findings challenge the conventional assumption of stable model maturation by demonstrating that intermediate checkpoints can exhibit superior reasoning and alignment capabilities compared to the final model. Furthermore, we show that strategic data selection effectively stabilizes the generalization process throughout pretraining.
This study addresses the challenges of indeterminate skill bottleneck resolution order and inefficient data mixing in large language model training. We propose a staged training framework integrating small proxy model exploration with a LogFloor closed-loop controller. This approach transforms bottleneck resolution trajectories into transferable curriculum learning structures, employing a "small-model reconnaissance and path transfer" mechanism to guide large models through sequential bottleneck breakthroughs. Experiments on Qwen2.5 demonstrate that this strategy reduces training tokens by an average of 56.2% and achieves approximately 39% computational savings through cross-scale transfer. These results indicate significant improvements in data efficiency for skill acquisition in large language models, offering a scalable solution to optimize training dynamics and resource utilization.
This study addresses the limitation that transferring capabilities from expert models to general-purpose language models typically relies on training or alignment. To overcome this, we propose a training-free heterogeneous model merging method that eliminates the need for gradient updates and semantic alignment. By leveraging parameter projection and interpolation, our approach directly facilitates cross-role knowledge transfer at the parameter level. Furthermore, we design two core merging strategies: Intersection-Merge and Activate-Prune-Merge. Experimental results demonstrate that the proposed method significantly enhances the performance of general-purpose models across embedding, reranking, and code generation tasks. This work establishes a novel paradigm for the efficient integration of heterogeneous models.
研究通过系统性调整学习率、批次大小等参数,对比LoRA和全微调方法,在不同模型和数据集上优化监督微调效果。