🤖 AI Summary
This work proposes three key techniques to enhance computational efficiency in large language model training while preserving or even improving performance. Selective Ground Truth training (SGT) achieves a 67% reduction in loss using only 15% of the original supervision signal. A depth-compression strategy combining layer averaging with unrolling attains the performance of larger models under a 2.5× parameter compression ratio. Additionally, a Mixture-of-Efficient-Experts (MoEE) architecture integrates dual experts to lower validation loss to 2.789. Together, these methods enable the construction of a highly efficient and high-performing Korean foundation model that significantly reduces computational overhead without compromising—indeed, sometimes enhancing—model efficacy.
📝 Abstract
We study three complementary techniques for training compute-efficient language models.
(1) Selective supervision and per-token efficiency. Selective Ground Truth Token Training (SGT) concentrates supervision on the ~15% of output tokens that carry semantic payload. Through positive gradient coupling in position-shared transformer weights -- a token-level instance of auxiliary-task transfer -- the remaining 85% of unsupervised tokens still improve substantially, giving a 4.5x per-supervised-token efficiency (at the step-100 eval optimum, ~67% of the full-sequence loss reduction is recovered from 15% of the supervision). We prove that this improvement on unsupervised tokens is guaranteed whenever the gradient coupling coefficient gamma-bar = 0.72 is positive (Theorem 1), and show the effect is a property of natural-language structure: it collapses on shuffled text.
(2) Depth compression with recurrent recovery. A 48-layer, 1B-parameter transformer is compressed to 6 layers (227M) by averaging adjacent layers and restored through learned recurrent unrolling. With 34 effective recurrent layers it reaches a held-out loss of 2.934, within measurement noise of a 566M dense model at 2.926 -- a 2.5x reduction in parameters.
(3) Fusion of compressed experts. Assembling several compressed models as a Mixture of Efficient Experts (MoEE) with multi-token prediction improves over each single expert at comparable active parameters: a 2-expert MoEE reaches loss 2.789 versus 2.926 for the best single compressed model.
We validate these techniques on CHERRY-1.8B, a Korean foundation model whose every trainable parameter derives from our own training runs. We are explicit throughout about the scope of the evidence (one model family, Korean data, loss-based metrics) and about which claims are established versus prospective.