Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing

📅 2026-08-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the impact of normalization placement—pre-normalization versus post-normalization—on large language model performance under curriculum-based depth growth training. By leveraging knowledge distillation, module freezing, and boundary representation diagnostics, the authors incrementally expand a student architecture under the guidance of a Qwen3-8B teacher model. Their analysis reveals, for the first time, a coupling effect between normalization placement and curriculum strategy: post-normalization consistently outperforms pre-normalization during depth expansion, achieving a cross-entropy reduction of 0.0328. This advantage stems from more stable boundary scales in the residual pathways following the addition of new modules, highlighting post-normalization’s superior adaptability in dynamically expanding architectures.
📝 Abstract
Pre-norm is the standard normalization placement in modern Transformers because it facilitates joint optimization of full-depth models. We ask whether this preference persists when depth is introduced through a curriculum. In curriculum depth growth, each appended block receives the boundary representation produced by a trained prefix, making normalization placement relevant to forward conditioning. We therefore test whether placement and training curriculum interact. In a controlled distillation study with a Qwen3-8B teacher and a nine-layer student, pre-norm and post-norm are indistinguishable under joint training, differing by $0.0004$ validation CE, while post-norm improves over pre-norm by $0.0328$ under curriculum growth, an order of magnitude larger. A post-joint control matched by student active-layer tokens remains worse than post-grow, which rules out compute as the sole explanation. The ranking crosses over during the curriculum: post-norm takes the lead once blocks are appended. Single-block and freeze controls localize the ranking change to block appending rather than shallow-block quality or retraining. Boundary diagnostics associate post-norm with stable residual scales and pre-norm with structural-token scale drift; on a fixed batch, the final pre-grow block is also nearly identity-mapped. Together with the phase-wise crossover, these observations are consistent with boundary-scale conditioning after new blocks are appended. The results motivate treating normalization placement and training curriculum as coupled design choices in this distillation setting.
Problem

Research questions and friction points this paper is trying to address.

normalization placement
curriculum depth growing
large language models
Transformer architecture
training curriculum
Innovation

Methods, ideas, or system contributions that make the work stand out.

normalization placement
curriculum depth growing
post-norm
Transformer distillation
boundary conditioning
🔎 Similar Papers
S
Sheng Ren
Nanjing University of Aeronautics and Astronautics, Nanjing, Jiangsu, China
Y
Yadong Wang
Nanjing University of Aeronautics and Astronautics, Nanjing, Jiangsu, China
N
Naiqiang Tan
Didichuxing Co. Ltd, Beijing, China
J
Jiangang Kong
Didichuxing Co. Ltd, Beijing, China
J
Jun Fang
Didichuxing Co. Ltd, Beijing, China
Rui Liu
Rui Liu
Didichuxing Co. Ltd, Beijing, China
J
Jun Wang
Didichuxing Co. Ltd, Beijing, China
K
Kai Chen
Didichuxing Co. Ltd, Beijing, China
L
Lipeng Liang
Didichuxing Co. Ltd, Beijing, China
Xiang Chen
Xiang Chen
Nanjing University of Science and Technology
Computer VisionImage ProcessingArtificial IntelligenceDeep Learning