🤖 AI Summary
To address the scalability bottlenecks of large language models arising from ever-increasing parameter and data requirements, this paper proposes an efficient decoder-only Transformer architecture. Methodologically, it employs a byte-level BPE tokenizer with a 128K vocabulary, integrated with Rotary Position Embeddings (RoPE), Grouped-Query Attention (GQA), RMSNorm, and SwiGLU activation—collectively reducing architectural redundancy and data dependency. Our key contribution is a lightweight yet highly effective model: trained on only 100 billion tokens with just 650 million parameters, it achieves 90% of the performance of a 1-billion-parameter baseline—reducing parameter count by 53% and training token volume by an order of magnitude. This demonstrates that careful architectural optimization enables a compelling trade-off between computational efficiency and modeling capability, without sacrificing downstream effectiveness.
📝 Abstract
We present Supernova, a 650M-parameter decoder-only transformer that demonstrates how careful architectural design and tokenization innovation can achieve the performance of larger models while maintaining computational efficiency. Our architecture combines Rotary Positional Embeddings (RoPE), Grouped Query Attention (GQA) with a 3:1 compression ratio, RMSNorm for computational efficiency, and SwiGLU activation functions. A critical innovation is our custom 128,000-vocabulary byte-level BPE tokenizer, which achieves state-of-the-art compression performance. Through detailed analysis, we show that Supernova achieves 90% of the performance of 1B-parameter models while using 53% fewer parameters and requiring only 100B training tokens--an order of magnitude less than competing models. Our findings challenge the prevailing scaling paradigm, demonstrating that architectural efficiency and tokenization quality can compensate for reduced parameter counts.