Auditing Agent Actions through Query-Conditioned Attribution

πŸ“… 2026-09-27
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the high cost and reliance on complete trajectories inherent in specific query attribution for auditing LLM agents. To this end, it pioneers a query-conditioned action attribution paradigm and introduces the AΒ³Bench benchmark. Methodologically, the proposed approach leverages lightweight open-source models to integrate gradient saliency with semantic relevance for evidence ranking, thereby achieving low-cost, highly specific traceability without expensive perturbation or external large model intervention. Experimental results demonstrate improvements of 40.9% and 42.1% in source MRR and evidence MAP, respectively. Furthermore, the method surpasses frontier model baselines in end-to-end accuracy while reducing deployment latency by 29.9%.
πŸ“ Abstract
LLM agents increasingly take consequential actions through interactions with users, policies, and external tools. Auditing these agents requires automated attribution of realized actions to their historical basis. However, existing attribution formulations do not provide question-specific traces for diverse auditing objectives. Additionally, when access to the acting model is limited (e.g., in API-only deployments), applicable methods commonly rely on costly input perturbations or external LLM analysis of complete trajectories. We therefore formulate $\textit{query-conditioned agent action attribution}, a new task that takes a natural-language auditing query as input and recovers the source and ordered intermediate evidence for the query-specified aspect of an action. We instantiate this task with $A^3Bench$, a benchmark comprising 1,396 auditing queries across policy basis, parameter provenance, failure propagation, and unsafe-behavior tracing. To enable efficient, query-specific attribution, we use small open-weight models as attribution proposers that combine query-conditioned gradient saliency with query-semantic relevance to rank history units. Our proposer consistently achieves stronger source and evidence rankings at lower inference cost than open-weight baselines, improving source MRR by up to 40.9\% and evidence MAP by 42.1\% with only two forward passes and one backward pass. Controlled evaluations confirm that our proposer improves attribution specificity by adapting its rankings to fine-grained changes in the auditing query. Building on a proposer ensemble, our end-to-end system surpasses the strongest frontier-model baseline in source accuracy (64.5\% vs.\ 60.4\%) while reducing empirical deployment latency by 29.9\% relative to the fastest frontier API baseline. Code and data will be released after the initial review period following final validation and cleanup.
Problem

Research questions and friction points this paper is trying to address.

Agent action attribution
LLM agents
Auditing
Query-conditioned
Model access limitation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Query-Conditioned Attribution
LLM Agent Auditing
Gradient Saliency
Attribution Proposer
A3Bench
πŸ”Ž Similar Papers
No similar papers found.