CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield

📅 2026-06-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work proposes three key techniques to enhance computational efficiency in large language model training while preserving or even improving performance. Selective Ground Truth training (SGT) achieves a 67% reduction in loss using only 15% of the original supervision signal. A depth-compression strategy combining layer averaging with unrolling attains the performance of larger models under a 2.5× parameter compression ratio. Additionally, a Mixture-of-Efficient-Experts (MoEE) architecture integrates dual experts to lower validation loss to 2.789. Together, these methods enable the construction of a highly efficient and high-performing Korean foundation model that significantly reduces computational overhead without compromising—indeed, sometimes enhancing—model efficacy.
📝 Abstract
We study three complementary techniques for training compute-efficient language models. (1) Selective supervision and per-token efficiency. Selective Ground Truth Token Training (SGT) concentrates supervision on the ~15% of output tokens that carry semantic payload. Through positive gradient coupling in position-shared transformer weights -- a token-level instance of auxiliary-task transfer -- the remaining 85% of unsupervised tokens still improve substantially, giving a 4.5x per-supervised-token efficiency (at the step-100 eval optimum, ~67% of the full-sequence loss reduction is recovered from 15% of the supervision). We prove that this improvement on unsupervised tokens is guaranteed whenever the gradient coupling coefficient gamma-bar = 0.72 is positive (Theorem 1), and show the effect is a property of natural-language structure: it collapses on shuffled text. (2) Depth compression with recurrent recovery. A 48-layer, 1B-parameter transformer is compressed to 6 layers (227M) by averaging adjacent layers and restored through learned recurrent unrolling. With 34 effective recurrent layers it reaches a held-out loss of 2.934, within measurement noise of a 566M dense model at 2.926 -- a 2.5x reduction in parameters. (3) Fusion of compressed experts. Assembling several compressed models as a Mixture of Efficient Experts (MoEE) with multi-token prediction improves over each single expert at comparable active parameters: a 2-expert MoEE reaches loss 2.789 versus 2.926 for the best single compressed model. We validate these techniques on CHERRY-1.8B, a Korean foundation model whose every trainable parameter derives from our own training runs. We are explicit throughout about the scope of the evidence (one model family, Korean data, loss-based metrics) and about which claims are established versus prospective.
Problem

Research questions and friction points this paper is trying to address.

compute-efficient language models
selective supervision
depth compression
Mixture of Experts
parameter efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Selective Ground Truth Token Training
Recurrent Unrolling
Depth Compression
Mixture of Efficient Experts
Gradient Coupling
D
Dohyeon Kwon
AX Architect, TeamSparta Inc., Seoul, South Korea
Y
Youngjin Park
TeamSparta Inc., Seoul, South Korea