Small Language Models for Agentic Systems: A Survey of Architectures, Capabilities, and Deployment Trade offs

📅 2025-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses the high cost, low efficiency, and unreliable outputs of large language models (LLMs) in agent systems by proposing a hybrid agent architecture that prioritizes small language models (SLMs) with dynamic LLM fallback. Methodologically, it introduces an uncertainty-aware routing mechanism and a validator cascade, integrated with guided decoding, JSON Schema–enforced structural constraints, type-safe function registration, and LoRA/QLoRA fine-tuning—leveraging efficient inference frameworks including vLLM, SGLang, and XGrammar. Key contributions include: (1) defining production-oriented evaluation metrics—e.g., cost per successful task and executable call rate; (2) matching or exceeding LLM performance on function-calling and RAG tasks; and (3) reducing inference latency by 10–100×, significantly lowering per-request energy consumption and token cost, thereby enabling on-device deployment.

Technology Category

Cognitive Modeling & Cognitive Systems: Agent ArchitecturesPlanning, Routing, and Scheduling: Planning with Language ModelsMultiagent Systems: Agent Communication

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Agentic search
📝 Abstract
Small language models (SLMs; 1-12B params, sometimes up to 20B) are sufficient and often superior for agentic workloads where the objective is schema- and API-constrained accuracy rather than open-ended generation. We synthesize recent evidence across open and proprietary SLMs (Phi-4-Mini, Qwen-2.5-7B, Gemma-2-9B, Llama-3.2-1B/3B, Ministral-3B/8B, Apple on-device 3B, DeepSeek-R1-Distill) and connect it to modern evaluations (BFCL v3/v4, StableToolBench) and serving stacks (vLLM, SGLang, TensorRT-LLM) paired with guided decoding libraries (XGrammar, Outlines). We formalize SLM-default, LLM-fallback systems with uncertainty-aware routing and verifier cascades, and propose engineering metrics that reflect real production goals: cost per successful task (CPS), schema validity rate, executable call rate, p50/p95 latency, and energy per request. Guided decoding, strict JSON Schema outputs, and validator-first tool execution close much of the capability gap with larger models and often let SLMs match or surpass LLMs on tool use, function calling, and RAG at 10x-100x lower token cost with materially better latency and energy. We provide design patterns for agent stacks that prioritize SLMs: schema-first prompting, type-safe function registries, confidence scoring with verifier rollups, and lightweight adaptation via LoRA/QLoRA. We also delineate limits where fallback remains valuable (open-domain reasoning and some long-horizon planning). The result is a practical blueprint for building fast, inexpensive, and reliable agents that default to SLMs while preserving headroom with targeted LLM assistance. Keywords: small language models, agents, function calling, structured outputs, JSON Schema, guided decoding, LoRA/QLoRA, routing, energy efficiency, edge inference
Problem

Research questions and friction points this paper is trying to address.

Optimizing small language models for agentic systems with constrained accuracy
Developing SLM-default systems with uncertainty-aware routing and verifier cascades
Closing capability gaps with larger models through guided decoding techniques
Innovation

Methods, ideas, or system contributions that make the work stand out.

SLMs with guided decoding for structured outputs
Uncertainty-aware routing with verifier cascades
Schema-first prompting and lightweight adaptation techniques
🔎 Similar Papers
2023-08-22Frontiers Comput. Sci.Citations: 866
R
Raghav Sharma
Northeastern University, Boston, USA Atlanta, USA
M
Manan Mehta
University of Southern California, USA New York, USA