Koa-action: Fast and Consistent Structured Decision Making with Generative LLMs

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high generation latency and verbose outputs of large language models in atomic decision-making tasks such as classification by proposing the Koa-action framework. This framework introduces atomic label tokens and integrates supervised fine-tuning with constrained decoding to transform multi-step generation into deterministic single-token output, enabling ultra-fast inference while preserving multimodal capabilities. Experimental results demonstrate that the proposed method achieves an intent routing accuracy of 85.5% and accelerates inference speed several-fold compared to state-of-the-art models. By delivering high accuracy alongside extremely low and stable latency, this work effectively balances inference efficiency with task flexibility.
📝 Abstract
Industry applications often demand low-latency classification, yet current large language model (LLM) approaches remain poorly suited for latency-critical applications. Existing prompting and constrained decoding produce verbose, multi-token outputs that require expensive token-by-token generation, while encoder-based models achieve faster inference but sacrifice task flexibility. We propose Koa-action, a framework for low-latency atomic actions -- fast, single-step decisions such as classification, semantic endpointing, Boolean checks, and scoring -- formulated as constrained generation with single-token outputs. By introducing atomic label tokens and applying supervised fine-tuning, our method reduces classification to a deterministic one-step decoding problem. Across standard benchmarks, Koa-action delivers competitive accuracy with consistently low and stable latency. On a production intent-routing benchmark, Koa-action reaches 85.5% accuracy -- competitive with the strongest frontier models (Claude-4.8-Opus, Gemini-Pro-3.1) and ahead of GPT-5 and Gemini-2.5-Pro -- while answering in about half a second, several-fold faster than every frontier model (up to ~7.5x at the median) under identical serving conditions. Against the dedicated single-token system Jev/TypeSafe, Koa-action is competitive on accuracy and faster at the median, while also handling multimodal inputs and multi-label outputs that single-label text systems do not.
Problem

Research questions and friction points this paper is trying to address.

low-latency classification
large language models
atomic actions
constrained generation
single-token decoding
Innovation

Methods, ideas, or system contributions that make the work stand out.

low-latency inference
single-token generation
constrained decoding
supervised fine-tuning
atomic label tokens
S
Shenghong Dai
Salesforce AI, University of Wisconsin–Madison
S
Shiva Kumar Pentyala
Salesforce AI
Y
Yingchi Liu
Salesforce AI
S
Shubham Mehrotra
Salesforce AI
Suman Banerjee
Suman Banerjee
Department of CSE, IIT Jammu
Algorithmic Data ManagementSocial Network AnalysisGraph Theory and Graph AlgorithmsParameterized Complexity
J
James Zhu
Salesforce AI
B
Bin Bi
Salesforce AI
Sitaram Asur
Sitaram Asur
Director, Salesforce (formerly in HP Labs)
Machine LearningNLPSocial Networks
P
Phil Mui
Salesforce AI