🤖 AI Summary
This work addresses the vulnerability of large language model (LLM) agents to unauthorized evidence in mixed-trust contexts, which can lead to policy-violating decisions. The authors propose a goal-oriented, fine-grained authorization auditing framework that isolates the influence of evidence source authority by fixing task specifications, propositions, stances, and policies, while systematically varying only the trustworthiness of evidence sources. Through contextual subset ablation and multi-model comparison, they quantify how unauthorized information alters action selection. The study introduces a novel auditing mechanism that separately annotates contextual factors with respect to tool-use and parameter-setting goals, alongside a controlled stress-testing protocol. In 450 controlled tasks, 5.4% of actions changed due to source differences, with 2.4% retaining conflicting unauthorized evidence—indicating that while LLMs perceive source cues, they remain susceptible to their influence.
📝 Abstract
LLM agents choose tools and arguments from context that mixes user requests, tool outputs, retrieved records, memory, and untrusted text. Evidence can be relevant without being authorized to determine a decision, so a correct action need not be grounded only in permitted evidence. We introduce a target-specific authorization audit that labels context factors separately for each tool and argument target. Its primary test holds the task, proposition, position, and policy fixed while changing only the proposition's source authority. We then test behavior when valid evidence is weakened and use context-subset interactions as a secondary localization diagnostic. Across 450 controlled next-action tasks and multiple open-weight LLM families, trusted and untrusted variants produce different actions in 5.4 percent of competing cases versus 1.7 percent of supporting cases. Under controlled degradation, unauthorized competition is retained in a full-correct, mixed-error, clean-correct pattern in 2.4 percent of comparisons, with a 95 percent confidence interval from 2.1 to 3.0 percent. These are controlled stress-set rates, not deployment prevalence. The models respond to textual source-authority cues, but this does not prevent untrusted evidence from influencing their actions.