🤖 AI Summary
Small language models (SLMs) face a fundamental trade-off between efficiency and performance when deployed on edge devices. This paper systematically surveys SLM architecture design, training paradigms, and lightweighting techniques, proposing the first taxonomy for SLM optimization tailored to mobile and edge platforms. We introduce a unified, open-source benchmark suite covering model pruning, quantization, knowledge distillation, and structured compression—enabling reproducible, multi-dimensional evaluation of accuracy–efficiency trade-offs. Our framework constitutes the first comprehensive, open, and reproducible SLM assessment infrastructure. It clarifies current Pareto bottlenecks in the accuracy–latency–memory–energy space and identifies promising future directions, particularly hardware-aware adaptation and synergistic compression strategies. The work provides both theoretical foundations and practical blueprints for developing efficient, compact language models suitable for resource-constrained environments.
📝 Abstract
Small Language Models (SLMs) have gained substantial attention due to their ability to execute diverse language tasks successfully while using fewer computer resources. These models are particularly ideal for deployment in limited environments, such as mobile devices, on-device processing, and edge systems. In this study, we present a complete assessment of SLMs, focussing on their design frameworks, training approaches, and techniques for lowering model size and complexity. We offer a novel classification system to organize the optimization approaches applied for SLMs, encompassing strategies like pruning, quantization, and model compression. Furthermore, we assemble SLM's studies of evaluation suite with some existing datasets, establishing a rigorous platform for measuring SLM capabilities. Alongside this, we discuss the important difficulties that remain unresolved in this sector, including trade-offs between efficiency and performance, and we suggest directions for future study. We anticipate this study to serve as a beneficial guide for researchers and practitioners who aim to construct compact, efficient, and high-performing language models.