small language model deployment

Designs and implements deployment systems for compact (small) language models, including model compression (quantization, pruning, distillation), runtime and serving optimizations, and packaging for resource‑constrained hardware to achieve low‑latency, low‑energy inference. Builds and evaluates SLM‑based optimization and ensemble‑inference strategies to trade off accuracy, robustness, and computational cost, and analyzes model performance, memory use, and energy characteristics in target runtime environments.

smalllanguagemodeldeployment

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.41
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Small Language Models: Survey, Measurements, and Insights

Sep 24, 2024
ZL
Zhenyan Lu
🏛️ Beijing University of Posts and Telecommunications | Peng Cheng Laboratory | Helixon Research | University of Cambridge

Small language models (SLMs) remain underexplored academically, with no systematic evaluation framework addressing their architectural diversity, training data, optimization strategies, and on-device performance. Method: This work introduces the first multidimensional benchmarking suite covering model architecture, pretraining data, training methodology, and edge-device efficiency, empirically evaluating 70 open-source SLMs (100M–5B parameters) across commonsense reasoning (MMLU), mathematical reasoning (GSM8K), code generation (HumanEval), in-context learning, and real-world edge deployment metrics—including latency and memory footprint. Contribution/Results: We uncover nonlinear trade-offs between capability and efficiency; identify SLMs approaching large language model (LLM) performance on specific tasks; reveal that parameter count and data quality exhibit diminishing returns beyond certain thresholds; and release the first reproducible, standardized benchmark dataset for on-device SLM inference—enabling rigorous, comparable research toward democratized edge AI.

Analyze SLMs' architectures and trainingBenchmark SLMs' on-device performance costsStudy small language models' capabilities

Must-Read Papers

Most classic and influential ideas
View more

Small Language Models: Architectures, Techniques, Evaluation, Problems and Future Adaptation

May 26, 2025
TH
Tanjil Hasan Sakib
🏛️ American International University-Bangladesh | Cornell University

Small language models (SLMs) face a fundamental trade-off between efficiency and performance when deployed on edge devices. This paper systematically surveys SLM architecture design, training paradigms, and lightweighting techniques, proposing the first taxonomy for SLM optimization tailored to mobile and edge platforms. We introduce a unified, open-source benchmark suite covering model pruning, quantization, knowledge distillation, and structured compression—enabling reproducible, multi-dimensional evaluation of accuracy–efficiency trade-offs. Our framework constitutes the first comprehensive, open, and reproducible SLM assessment infrastructure. It clarifies current Pareto bottlenecks in the accuracy–latency–memory–energy space and identifies promising future directions, particularly hardware-aware adaptation and synergistic compression strategies. The work provides both theoretical foundations and practical blueprints for developing efficient, compact language models suitable for resource-constrained environments.

Addressing efficiency-performance trade-offs in SLM optimization.Evaluating design and training of Small Language Models (SLMs).Proposing future directions for compact, high-performing SLMs.

Small language models (SLMs) suffer from low quantization efficiency on edge devices, and existing large language model (LLM) quantization methods transfer poorly to SLMs. Method: We introduce SLMQuant—the first systematic quantization benchmark for SLMs—covering diverse architectures, tasks, and quantization strategies (weight-only, activation-aware, and mixed-precision). Through empirical analysis, we identify SLMs’ higher quantization sensitivity and distinct bottlenecks compared to LLMs, leading to a set of SLM-specific compression design principles. Contribution/Results: Directly applying LLM quantization schemes severely degrades SLM performance. SLMQuant fills a critical evaluation gap and provides reproducible, task-aware optimization guidance. Experiments show it improves edge deployment efficiency by up to 2.3× while preserving accuracy, establishing a principled foundation for efficient SLM deployment in resource-constrained environments.

Establishing tailored compression principles for efficient SLM deploymentEvaluating quantization effectiveness for Small Language Models on edge devicesIdentifying unique quantization bottlenecks in SLMs versus LLMs

Edge-First Language Model Inference: Models, Metrics, and Tradeoffs

May 22, 2025
SJ
SiYoung Jang
🏛️ Nokia Bell Labs | EURECOM

This study addresses efficient inference deployment of small language models (SLMs) in edge–cloud continuum environments, balancing low latency, strong privacy, low operational cost, and reliability. We propose a platform-level adaptive inference paradigm grounded in an edge-first design principle and a quantitative trade-off framework—rejecting one-size-fits-all strategies. Our methodology integrates model compression, edge device- and cluster-level benchmarking, and multi-dimensional evaluation across latency, cost, and reliability. Key contributions include: (i) the first systematic characterization of feasibility boundaries for SLMs on resource-constrained edge devices; (ii) a reusable, context-aware deployment decision guide; and (iii) empirical improvements over pure cloud-based inference—achieving 37% average latency reduction and 22% lower operational cost on representative tasks. The framework enables principled, environment-aware SLM deployment across heterogeneous edge–cloud infrastructures.

Balancing edge-cloud tradeoffs for adaptive LM inference systemsDeploying Language Models from cloud to edge for cost and latency benefitsEvaluating Small Language Models on resource-constrained edge platforms

Towards Pareto Optimal Throughput in Small Language Model Serving

Apr 04, 2024
PG
Pol G. Recasens
🏛️ Barcelona Supercomputing Center | IBM Research | Universitat Politècnica de Catalunya

This work addresses throughput optimization for small language model (SLM) inference under resource-constrained settings. Methodologically, it establishes, for the first time, that SLMs—owing to their substantial memory efficiency—can achieve near-theoretical-peak throughput on a single accelerator; further, it systematically demonstrates that model replication is the key strategy for improving hardware utilization and energy efficiency. Integrating system-level performance–energy co-benchmarking, memory bandwidth modeling, and throughput–latency trade-off analysis, the paper develops an interpretable optimization framework. Empirical evaluation shows that the proposed paradigm achieves Pareto-optimal throughput for SLM serving on a single GPU, improving hardware utilization by up to 3.2× and energy efficiency by up to 2.8×, thereby offering a novel pathway for efficient deployment of lightweight large language models.

Achieve Pareto-optimal throughput with SLMsBenchmark SLM inference performance and energy levelsImprove resource utilization via model replication

Fine-Tuning and Deploying Large Language Models Over Edges: Issues and Approaches

Aug 20, 2024
YD
Yanjie Dong
🏛️ Artificial Intelligence Research Institute | Guangdong-Hong Kong-Macao Joint Laboratory for Emotional Intelligence and Pervasive Computing | Shenzhen MSU-BIT University | School of Medical Technology | Beijing Institute of Technology

Edge-device GPUs face severe memory constraints, hindering fine-tuning and multimodal extension of large language models (LLMs). Method: We systematically survey memory-efficient fine-tuning techniques (e.g., LoRA, QLoRA, Adapters) and model compression methods (e.g., quantization, pruning, knowledge distillation, sparse training), and propose, for the first time, a synergistic fine-tuning-and-compression paradigm tailored for edge deployment. We design a unified evaluation framework that quantifies trade-offs across three dimensions: energy efficiency, hardware compatibility, and multimodal generalization capability. Contribution/Results: We establish the first taxonomy of LLM lightweighting techniques specifically for edge deployment, characterizing each method’s performance in GPU memory footprint, inference latency, accuracy retention, and cross-platform adaptability. Our work provides both theoretical foundations and practical guidelines for sustainable on-device AI deployment.

Deploying large-scale multi-modal foundation models at network edgesEfficient fine-tuning of LLMs on edge devices with limited memoryReducing operational costs for LLM deployment via compression techniques

Latest Papers

What's happening recently
View more

Deploying large code language models on resource-constrained local devices remains challenging due to hardware limitations, which compromises privacy preservation, inference latency, and offline availability. This work proposes Ditto, a method that co-optimizes model compression and inference program generation to compile large code models into lightweight executables suitable for statically typed languages such as C. Ditto innovatively integrates bounded-error product quantization with quantization-aware inference program synthesis and extends the LLVM compiler to automatically replace general matrix-vector (GEMV) operations with efficient BLAS calls. Evaluated across three prominent code language models, Ditto achieves up to 10.5× speedup and 6.4× reduction in memory footprint, with an average drop of only 0.27% in pass@1 accuracy.

Code LLMslocal deploymentmodel efficiency

Deploying large language models on resource-constrained hardware faces significant challenges, including the trade-off between accuracy and efficiency and the complexity of quantization hyperparameter tuning. This work proposes the Hardware-Aware Quantization Agent (HAQA), which leverages a large language model to automatically optimize quantization hyperparameters and adapt them to target hardware, enabling cross-platform adaptive quantization strategies. By doing so, HAQA substantially reduces the need for manual intervention and streamlines the deployment pipeline. Experiments on the Llama family of models demonstrate that HAQA achieves up to 2.3× speedup in inference latency and throughput while maintaining or even improving model accuracy, outperforming conventional non-optimized deployment approaches.

accuracydeploymenthardware constraints

This work addresses the high energy consumption of large language model (LLM) inference and the inefficiency of conventional parameter tuning methods, which often require days of computation and struggle to adapt to diverse hardware and system constraints. The authors propose a human-in-the-loop optimization framework that integrates conversational LLMs with human feedback, leveraging an enhanced prompt template to enable rapid, adaptive search over inference runtime parameters. Their approach significantly improves tuning efficiency, converging to configurations below a target energy threshold in just 3.4 prompts on average—outperforming baseline methods such as Sobol sampling, which requires 5.2 prompts—and consistently achieves lower energy consumption per token, demonstrating clear advantages in both convergence speed and energy efficiency.

Energy EfficiencyLarge Language ModelsModel Inference

This study addresses the significant yet underexplored impact of system-level design on energy consumption in large language model (LLM) inference. Through empirical evaluation on an NVIDIA H100 platform, we systematically analyze how numerical precision, dynamic batching, and request scheduling—implemented via quantization, dynamic batching, and Hugging Face TGI—affect inference energy efficiency and latency. Our findings reveal that low-precision computation reduces energy only during compute-intensive phases, dynamic batching substantially improves decoding energy efficiency, and structured request scheduling can reduce per-request energy consumption by up to two orders of magnitude. Building on these insights, we propose a phase-aware energy efficiency analysis framework that offers both theoretical grounding and practical guidance for optimizing LLM inference systems.

batchingenergy efficiencyLLM inference

Hot Scholars

MM

Michele Magno

ETH Zurich
Wireless sensor networksSmart Sensors and Internet of ThingsWake up RadioPower management
SR

Shashi Raj Pandey

Aalborg University
Network EconomicsGame TheoryWireless NetworksDistributed Learning
AS

Agam Shah

PhD Candidate, Georgia Institute of Technology
Natural Language ProcessingFinanceData ScienceComputational Science
PP

Petar Popovski

Professor, Connectivity, Aalborg University, Denmark
Communication TheoryWireless Communications5G6G
HB

Honglin Bao

University of Chicago
InnovationAINLPSocial Networks