G-CARB: Graph-Localized Conformal Agent Risk Budget for Compositional Harm

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the compositional hazards encountered by small language model agents during tool calling, where existing safety monitoring incurs substantial overhead and complicates consequence tracing. To mitigate these challenges, this work proposes a graph-localized conformal risk budgeting framework that innovatively integrates conformal prediction with graph dependency analysis. By recording execution histories and filtering scoring evidence along observable dependencies, the framework achieves lightweight risk assessment and safe stopping control without requiring additional inference. Evaluated on the AgentDojo benchmark, the proposed approach significantly reduces input logging volume while improving autonomous task completion rates. Furthermore, it effectively prevents unwarranted termination of benign operations, providing low-overhead protection against potential leakage risks across agent tool invocations.
📝 Abstract
Small language model (SLM) agents need safety controls that track consequences across tool calls with little monitoring overhead. A private read, for example, becomes a leak when a later action sends that data outside the system. We introduce CARB (Conformal Agent Risk Budget), which calibrates when to stop an agent using a ledger of harm incurred before stopping. Under exchangeable episodes, standard conformal risk control bounds this declared loss in expectation over calibration and a future episode. G-CARB selects scorer evidence along observable dependencies from private sources to outgoing actions. The ledger still covers the entire executed history, and computing the gate score requires no additional language-model inference. On AgentDojo replay with two 14B backbones, G-CARB roughly halves scorer-input records at intermediate risk budgets while improving autonomous task completion relative to full-prefix scoring; random context of the same size achieves similar gains. Controlled examples show how retaining the relevant dependency can further avoid stopping benign work.
Problem

Research questions and friction points this paper is trying to address.

small language model agents
compositional harm
safety control
risk budget
tool calls
Innovation

Methods, ideas, or system contributions that make the work stand out.

Conformal Risk Control
Graph-Localized Scoring
Small Language Model Agents
Compositional Harm
Risk Budget
🔎 Similar Papers
No similar papers found.