Toward SLM-based agentic task-tool intent matching

๐Ÿ“… 2026-10-02
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the challenge that AI agent tool invocations often deviate from task intent, while conventional authorization mechanisms fail to verify underlying cognitive logic. To overcome this, we propose a novel paradigm for task-tool intent alignment verification leveraging small language models (SLMs). Specifically, we construct a multi-tool task dataset spanning multiple MCP servers and train an SLM through prompt optimization, supervised fine-tuning, and Group Relative Policy Optimization (GRPO) reinforcement learning, enabling it to function as a real-time relevance classifier that evaluates invocation-task matching. This approach achieves low-latency, automated intent consistency detection independent of execution environments, thereby significantly enhancing the safety and controllability of agent systems.
๐Ÿ“ Abstract
Tool-equipped AI agents use tool calls to access data and act on external systems. Horizontal growth of agentic systems increases the number of these interactions, and further motivates the need for automated, per-call oversight that can operate at low latency and/or on-prem. Conventional authorization schemes can determine whether an agent is allowed to invoke a tool, but cannot assess the agent's underlying cognition, specifically, whether the tool selection represents a logical, relevant step toward satisfying the intent of the task or not. Consequently, an allowed call may still deviate from the task's intent: a rogue agent might deviate the calls or nudge other agents to make a combination of calls that would not align with the intent of the task. Therefore, every call needs to be verified. In this study we investigate the applicability of Small Language Models (SLMs) to this purpose: an SLM functions as a task-tool relevance classifier that evaluates every selected tool independently against the assigned task and returns a relevance signal for downstream enforcement. Equipped with a novel dataset with multi-tool tasks whose required tools span distinct Model Context Protocol (MCP) servers, we used prompt-optimization, supervised fine-tuning, and reinforcement learning through GRPO to optimize and specialize SLMs.
Problem

Research questions and friction points this paper is trying to address.

AI agents
task-tool intent matching
tool call verification
Small Language Models
Model Context Protocol
Innovation

Methods, ideas, or system contributions that make the work stand out.

Small Language Models
Intent Matching
GRPO
Model Context Protocol
Agentic Oversight
๐Ÿ”Ž Similar Papers
C
Chiara Troiani
Cisco Systems, Switzerland
Arash Salarian
Arash Salarian
Cisco Systems, Switzerland
Majed El Helou
Majed El Helou
ETH Zurich
Computational ImagingImage ProcessingComputer VisionDeep Learning
Benjamin Ryder
Benjamin Ryder
ETH Zรผrich
Artificial IntelligenceMachine LearningDriving Data
J
Jean Diaconu
Cisco Systems, Switzerland
H
Hervรฉ Muyal
Cisco Systems, Switzerland
M
Marcelo Yannuzzi
Cisco Systems, Switzerland