POLAR: Ontology-Guided Risk Prevention for Tool-Calling LLM Agents

๐Ÿ“… 2026-10-06
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the absence of preemptive risk prevention mechanisms in LLM tool agents by proposing POLAR, a novel defense framework. POLAR introduces a dual-layer ontology-based reversibility grading score and a candidate inverse sequence derivation mechanism. By leveraging structured ontological modeling to assess operation reversibility, the framework prunes high-risk tool calls prior to execution and supports integration with small-model agents. Experimental evaluations on ฯ„ยฒ-bench demonstrate that POLAR effectively improves task rewards for specific scenarios while delineating the trade-off boundary between utility and safety. Ultimately, this work provides auditable, preemptive safety guardrails for LLM agents, mitigating risks before irreversible actions occur.
๐Ÿ“ Abstract
LLM tool-use agents operate in dynamic environments where many actions carry operational risk. However, most safety mechanisms react only after errors manifest. Existing pre-emptive approaches either fine-tune the agent on chain-of-thought deliberation or compile natural-language guardrails into runtime checks, but they do so without exposing a structural, auditable verdict. We propose POLAR, a guardrail framework for small tool-calling agents that assesses reversibility through a structured two-layer ontology. POLAR assigns each action a graded reversibility score by deriving a candidate inverse sequence; calls failing a threshold are pruned before execution. Evaluated on $ฯ„^2$-bench across six agent models, POLAR improves mean task reward by 0.11 to 0.18 points on airline for four of six agents, but only eight of eighteen model--domain cells improve overall; retail and stronger agents often regress. POLAR provides an auditable structural check and characterizes its task-utility trade-offs. Reward is not a direct measure of prevented harm.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
tool-calling
risk prevention
safety guardrails
action reversibility
Innovation

Methods, ideas, or system contributions that make the work stand out.

Ontology-Guided Guardrail
Action Reversibility
Tool-Calling Agents
Preemptive Risk Prevention
Auditable Safety
๐Ÿ”Ž Similar Papers