Supernova: Achieving More with Less in Transformer Architectures

📅 2025-07-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the scalability bottlenecks of large language models arising from ever-increasing parameter and data requirements, this paper proposes an efficient decoder-only Transformer architecture. Methodologically, it employs a byte-level BPE tokenizer with a 128K vocabulary, integrated with Rotary Position Embeddings (RoPE), Grouped-Query Attention (GQA), RMSNorm, and SwiGLU activation—collectively reducing architectural redundancy and data dependency. Our key contribution is a lightweight yet highly effective model: trained on only 100 billion tokens with just 650 million parameters, it achieves 90% of the performance of a 1-billion-parameter baseline—reducing parameter count by 53% and training token volume by an order of magnitude. This demonstrates that careful architectural optimization enables a compelling trade-off between computational efficiency and modeling capability, without sacrificing downstream effectiveness.

Technology Category

Machine Learning: Deep Neural Architectures and Foundation ModelsNatural Language Processing: (Large) Language ModelsComputer Vision: Large Vision Models

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Efficiency and scalability of Web search enginesEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
We present Supernova, a 650M-parameter decoder-only transformer that demonstrates how careful architectural design and tokenization innovation can achieve the performance of larger models while maintaining computational efficiency. Our architecture combines Rotary Positional Embeddings (RoPE), Grouped Query Attention (GQA) with a 3:1 compression ratio, RMSNorm for computational efficiency, and SwiGLU activation functions. A critical innovation is our custom 128,000-vocabulary byte-level BPE tokenizer, which achieves state-of-the-art compression performance. Through detailed analysis, we show that Supernova achieves 90% of the performance of 1B-parameter models while using 53% fewer parameters and requiring only 100B training tokens--an order of magnitude less than competing models. Our findings challenge the prevailing scaling paradigm, demonstrating that architectural efficiency and tokenization quality can compensate for reduced parameter counts.
Problem

Research questions and friction points this paper is trying to address.

Achieves performance of larger models with fewer parameters
Improves efficiency via architectural design and tokenization
Reduces training tokens needed compared to competing models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Combines RoPE, GQA, RMSNorm, SwiGLU for efficiency
Uses custom 128K BPE tokenizer for compression
Achieves 90% performance with 53% fewer parameters
🔎 Similar Papers
No similar papers found.
Ovidius University of Constanţa
A
Andrei-Valentin Tanase
Faculty of Mathematics and Computer Science, “Ovidius” University of Constanţa, Romania
Elena Pelican
Elena Pelican
Assoc. Prof., Ovidius University of Constanta
Computer VisionMachine LearningNumerical Linear Algebra