On Device Agentic Operation Caches -- Classifier-Centric NL-to-Action Generation

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high latency, privacy risks, and substantial costs associated with processing natural language into operational commands using cloud-based large models. To overcome these limitations, this work proposes an on-device agent architecture featuring a novel classifier-based operation caching mechanism. By reformulating generative tasks as a classification paradigm, the approach decouples natural language understanding from complex reasoning dependencies, thereby enabling the localized execution of high-frequency actions. Evaluated on an Excel formula generation task, the proposed method reduces total inference costs by 56% and decreases response latency fivefold upon cache hits. These results demonstrate that the framework provides an efficient new paradigm for edge-cloud collaborative interaction.
📝 Abstract
Agentic AI is increasingly being embedded in software applications to provide natural language interfaces to features and functionality. In most cases these agents are powered by enterprise (100+ billion parameter) or frontier class large language models that require substantial computational resources run and depend on cloud hosted inference to handle the task of transforming natural language inputs into actionable software operations. This reliance on cloud-hosted inference introduces substantial network latency on top of LLM inference times, creates data privacy concerns, and, given the costs of running these models, can rapidly escalate expenses associated with supporting agentic features. This paper introduces a novel means of converting the NL-to-Action problem from a generative one into a classification-centric formulation via on-device operation caches. These caches allow an agentic system to handle frequently occurring classes of actions completely on-device -- reducing latency, enhancing privacy, and lowering operational costs. We show that for a classic NL-to-Formula task, generating Excel Formula in response to user requests, this approach reduces total inference cost by 56% when compared to cloud-only model-routing based inference and, on cache hits, reduces the latency to response latency by 5x.
Problem

Research questions and friction points this paper is trying to address.

Agentic AI
Natural Language to Action
Cloud Inference
Latency
Data Privacy
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Device Operation Caches
Classifier-Centric
NL-to-Action
Agentic AI
Model Routing
🔎 Similar Papers
No similar papers found.