Long-Horizon Agent Trajectory Attribution: A Unified Benchmark and Fine-Grained Annotation Framework

📅 2026-08-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing benchmarks struggle to perform fine-grained behavioral attribution on long-horizon agent trajectories, making it difficult to identify the specific components responsible for particular outcomes. This work introduces the trajectory attribution task for the first time, proposing a unified component-based trajectory representation and a fine-grained annotation framework that supports two core evaluation objectives: attribution localization and attribution chain recovery. Leveraging AgentDojo and Agent3Sigma, the authors annotate over 1,300 trajectories covering diverse scenarios such as task alignment, unsafe behaviors, and safety refusals. Baseline methods are established through incremental contribution analysis and component-level leave-one-out perturbations. This study provides the first systematic and scalable benchmark for attribution research in long-horizon autonomous agents.
📝 Abstract
Large language model (LLM) agents increasingly operate through long-horizon trajectories involving user instructions, tool use, external observations, and memory. Existing benchmarks primarily evaluate behavioral outcomes but provide limited support for fine-grained attribution analysis. We introduce trajectory attribution and develop a benchmark and annotation framework for this task. The benchmark organizes heterogeneous trajectories under a unified component schema and provides annotations of the primary attribution component, together with attack and execution chains where applicable. Instantiating the benchmark with trajectories from AgentDojo and the Stage and Canary settings of Agent3Sigma yields more than 1,300 annotated trajectories covering task-aligned actions, unsafe actions, and safety refusals. The benchmark defines two evaluation tasks, primary attribution localization and attribution-chain recovery, and provides reference baselines based on incremental trajectory contribution and component-level leave-one-out perturbation. It captures diverse attribution settings, including local and long-range attribution as well as structured attribution chains. Reference baseline results exhibit substantial performance differences across these settings, providing an initial characterization of the benchmark's attribution challenges. Beyond this initial instantiation, we release a reusable annotation skill that enables trajectories generated by new agent models to be standardized, annotated, and evaluated under the same framework. Project resources and future releases are available at https://github.com/chenjing-2024/agent-trajectory-attribution.
Problem

Research questions and friction points this paper is trying to address.

trajectory attribution
long-horizon agent
fine-grained annotation
attribution analysis
LLM agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

trajectory attribution
long-horizon agent
fine-grained annotation
attribution chain
unified benchmark
🔎 Similar Papers
No similar papers found.