🤖 AI Summary
This study addresses the high generation latency and verbose outputs of large language models in atomic decision-making tasks such as classification by proposing the Koa-action framework. This framework introduces atomic label tokens and integrates supervised fine-tuning with constrained decoding to transform multi-step generation into deterministic single-token output, enabling ultra-fast inference while preserving multimodal capabilities. Experimental results demonstrate that the proposed method achieves an intent routing accuracy of 85.5% and accelerates inference speed several-fold compared to state-of-the-art models. By delivering high accuracy alongside extremely low and stable latency, this work effectively balances inference efficiency with task flexibility.
📝 Abstract
Industry applications often demand low-latency classification, yet current large language model (LLM) approaches remain poorly suited for latency-critical applications. Existing prompting and constrained decoding produce verbose, multi-token outputs that require expensive token-by-token generation, while encoder-based models achieve faster inference but sacrifice task flexibility. We propose Koa-action, a framework for low-latency atomic actions -- fast, single-step decisions such as classification, semantic endpointing, Boolean checks, and scoring -- formulated as constrained generation with single-token outputs. By introducing atomic label tokens and applying supervised fine-tuning, our method reduces classification to a deterministic one-step decoding problem. Across standard benchmarks, Koa-action delivers competitive accuracy with consistently low and stable latency. On a production intent-routing benchmark, Koa-action reaches 85.5% accuracy -- competitive with the strongest frontier models (Claude-4.8-Opus, Gemini-Pro-3.1) and ahead of GPT-5 and Gemini-2.5-Pro -- while answering in about half a second, several-fold faster than every frontier model (up to ~7.5x at the median) under identical serving conditions. Against the dedicated single-token system Jev/TypeSafe, Koa-action is competitive on accuracy and faster at the median, while also handling multimodal inputs and multi-label outputs that single-label text systems do not.