Score
Designs and implements deployment systems for compact (small) language models, including model compression (quantization, pruning, distillation), runtime and serving optimizations, and packaging for resource‑constrained hardware to achieve low‑latency, low‑energy inference. Builds and evaluates SLM‑based optimization and ensemble‑inference strategies to trade off accuracy, robustness, and computational cost, and analyzes model performance, memory use, and energy characteristics in target runtime environments.
This paper addresses the high cost, low efficiency, and unreliable outputs of large language models (LLMs) in agent systems by proposing a hybrid agent architecture that prioritizes small language models (SLMs) with dynamic LLM fallback. Methodologically, it introduces an uncertainty-aware routing mechanism and a validator cascade, integrated with guided decoding, JSON Schema–enforced structural constraints, type-safe function registration, and LoRA/QLoRA fine-tuning—leveraging efficient inference frameworks including vLLM, SGLang, and XGrammar. Key contributions include: (1) defining production-oriented evaluation metrics—e.g., cost per successful task and executable call rate; (2) matching or exceeding LLM performance on function-calling and RAG tasks; and (3) reducing inference latency by 10–100×, significantly lowering per-request energy consumption and token cost, thereby enabling on-device deployment.
Small language models (SLMs) remain underexplored academically, with no systematic evaluation framework addressing their architectural diversity, training data, optimization strategies, and on-device performance. Method: This work introduces the first multidimensional benchmarking suite covering model architecture, pretraining data, training methodology, and edge-device efficiency, empirically evaluating 70 open-source SLMs (100M–5B parameters) across commonsense reasoning (MMLU), mathematical reasoning (GSM8K), code generation (HumanEval), in-context learning, and real-world edge deployment metrics—including latency and memory footprint. Contribution/Results: We uncover nonlinear trade-offs between capability and efficiency; identify SLMs approaching large language model (LLM) performance on specific tasks; reveal that parameter count and data quality exhibit diminishing returns beyond certain thresholds; and release the first reproducible, standardized benchmark dataset for on-device SLM inference—enabling rigorous, comparable research toward democratized edge AI.
Small language models (SLMs) face a fundamental trade-off between efficiency and performance when deployed on edge devices. This paper systematically surveys SLM architecture design, training paradigms, and lightweighting techniques, proposing the first taxonomy for SLM optimization tailored to mobile and edge platforms. We introduce a unified, open-source benchmark suite covering model pruning, quantization, knowledge distillation, and structured compression—enabling reproducible, multi-dimensional evaluation of accuracy–efficiency trade-offs. Our framework constitutes the first comprehensive, open, and reproducible SLM assessment infrastructure. It clarifies current Pareto bottlenecks in the accuracy–latency–memory–energy space and identifies promising future directions, particularly hardware-aware adaptation and synergistic compression strategies. The work provides both theoretical foundations and practical blueprints for developing efficient, compact language models suitable for resource-constrained environments.
Small language models (SLMs) suffer from low quantization efficiency on edge devices, and existing large language model (LLM) quantization methods transfer poorly to SLMs. Method: We introduce SLMQuant—the first systematic quantization benchmark for SLMs—covering diverse architectures, tasks, and quantization strategies (weight-only, activation-aware, and mixed-precision). Through empirical analysis, we identify SLMs’ higher quantization sensitivity and distinct bottlenecks compared to LLMs, leading to a set of SLM-specific compression design principles. Contribution/Results: Directly applying LLM quantization schemes severely degrades SLM performance. SLMQuant fills a critical evaluation gap and provides reproducible, task-aware optimization guidance. Experiments show it improves edge deployment efficiency by up to 2.3× while preserving accuracy, establishing a principled foundation for efficient SLM deployment in resource-constrained environments.
This study addresses efficient inference deployment of small language models (SLMs) in edge–cloud continuum environments, balancing low latency, strong privacy, low operational cost, and reliability. We propose a platform-level adaptive inference paradigm grounded in an edge-first design principle and a quantitative trade-off framework—rejecting one-size-fits-all strategies. Our methodology integrates model compression, edge device- and cluster-level benchmarking, and multi-dimensional evaluation across latency, cost, and reliability. Key contributions include: (i) the first systematic characterization of feasibility boundaries for SLMs on resource-constrained edge devices; (ii) a reusable, context-aware deployment decision guide; and (iii) empirical improvements over pure cloud-based inference—achieving 37% average latency reduction and 22% lower operational cost on representative tasks. The framework enables principled, environment-aware SLM deployment across heterogeneous edge–cloud infrastructures.
This work addresses throughput optimization for small language model (SLM) inference under resource-constrained settings. Methodologically, it establishes, for the first time, that SLMs—owing to their substantial memory efficiency—can achieve near-theoretical-peak throughput on a single accelerator; further, it systematically demonstrates that model replication is the key strategy for improving hardware utilization and energy efficiency. Integrating system-level performance–energy co-benchmarking, memory bandwidth modeling, and throughput–latency trade-off analysis, the paper develops an interpretable optimization framework. Empirical evaluation shows that the proposed paradigm achieves Pareto-optimal throughput for SLM serving on a single GPU, improving hardware utilization by up to 3.2× and energy efficiency by up to 2.8×, thereby offering a novel pathway for efficient deployment of lightweight large language models.
Edge-device GPUs face severe memory constraints, hindering fine-tuning and multimodal extension of large language models (LLMs). Method: We systematically survey memory-efficient fine-tuning techniques (e.g., LoRA, QLoRA, Adapters) and model compression methods (e.g., quantization, pruning, knowledge distillation, sparse training), and propose, for the first time, a synergistic fine-tuning-and-compression paradigm tailored for edge deployment. We design a unified evaluation framework that quantifies trade-offs across three dimensions: energy efficiency, hardware compatibility, and multimodal generalization capability. Contribution/Results: We establish the first taxonomy of LLM lightweighting techniques specifically for edge deployment, characterizing each method’s performance in GPU memory footprint, inference latency, accuracy retention, and cross-platform adaptability. Our work provides both theoretical foundations and practical guidelines for sustainable on-device AI deployment.
Deploying large code language models on resource-constrained local devices remains challenging due to hardware limitations, which compromises privacy preservation, inference latency, and offline availability. This work proposes Ditto, a method that co-optimizes model compression and inference program generation to compile large code models into lightweight executables suitable for statically typed languages such as C. Ditto innovatively integrates bounded-error product quantization with quantization-aware inference program synthesis and extends the LLVM compiler to automatically replace general matrix-vector (GEMV) operations with efficient BLAS calls. Evaluated across three prominent code language models, Ditto achieves up to 10.5× speedup and 6.4× reduction in memory footprint, with an average drop of only 0.27% in pass@1 accuracy.
Deploying large language models on resource-constrained hardware faces significant challenges, including the trade-off between accuracy and efficiency and the complexity of quantization hyperparameter tuning. This work proposes the Hardware-Aware Quantization Agent (HAQA), which leverages a large language model to automatically optimize quantization hyperparameters and adapt them to target hardware, enabling cross-platform adaptive quantization strategies. By doing so, HAQA substantially reduces the need for manual intervention and streamlines the deployment pipeline. Experiments on the Llama family of models demonstrate that HAQA achieves up to 2.3× speedup in inference latency and throughput while maintaining or even improving model accuracy, outperforming conventional non-optimized deployment approaches.
This work addresses the high energy consumption of large language model (LLM) inference and the inefficiency of conventional parameter tuning methods, which often require days of computation and struggle to adapt to diverse hardware and system constraints. The authors propose a human-in-the-loop optimization framework that integrates conversational LLMs with human feedback, leveraging an enhanced prompt template to enable rapid, adaptive search over inference runtime parameters. Their approach significantly improves tuning efficiency, converging to configurations below a target energy threshold in just 3.4 prompts on average—outperforming baseline methods such as Sobol sampling, which requires 5.2 prompts—and consistently achieves lower energy consumption per token, demonstrating clear advantages in both convergence speed and energy efficiency.
This study addresses the significant yet underexplored impact of system-level design on energy consumption in large language model (LLM) inference. Through empirical evaluation on an NVIDIA H100 platform, we systematically analyze how numerical precision, dynamic batching, and request scheduling—implemented via quantization, dynamic batching, and Hugging Face TGI—affect inference energy efficiency and latency. Our findings reveal that low-precision computation reduces energy only during compute-intensive phases, dynamic batching substantially improves decoding energy efficiency, and structured request scheduling can reduce per-request energy consumption by up to two orders of magnitude. Building on these insights, we propose a phase-aware energy efficiency analysis framework that offers both theoretical grounding and practical guidance for optimizing LLM inference systems.