🤖 AI Summary
This study addresses the prevalence of undesirable behaviors in agent deployment and the lack of real-world trajectory benchmarks by proposing an automated benchmark construction and evaluation system. The system extracts specific undesirable behaviors from authentic deployment trajectories and introduces a novel anchor verification mechanism coupled with a synthesis loop. By integrating programmable retrieval, large language model verification, and decision-point continuation techniques, it enables customized evaluation without environment replay while supporting the generation of user-defined behavioral specifications. Experimentally, 107 benchmarks comprising 4,125 instances were constructed, covering 95.3% of testing requirements. Notably, state-of-the-art models achieved an average pass rate of only 26.7%, profoundly revealing critical behavioral deficiencies in current AI agents.
📝 Abstract
An agent can complete a task while exhibiting undesirable behavior during execution. Developers need tests for the specific behaviors encountered in deployment, beyond fixed benchmark suites. We present TraceDance, an agent system that constructs targeted benchmarks from deployment traces for user-specified undesirable behaviors. For efficient construction, Anchor-and-Confirm combines programmable retrieval with candidate-level confirmation by a Flash large language model (LLM), while the Anchor Synthesis Loop generates and revises specifications for custom behaviors. The benchmarks use decision-point continuation to evaluate an LLM's next turn at a recorded decision point with a behavior-specific rubric, without a reference answer or environment replay. Experiments in coding and general tool use draw on 252,557 sessions and produce 107 benchmarks with 4,125 instances, fulfilling 95.3% of build-target requests. Both human annotators confirm the requested behavior in 84% of sampled instances, and the automated grader's agreement with human pass/fail judgments is comparable to that between the annotators. Nine frontier LLMs achieve a mean pass rate of only 26.7%, showing that they still struggle to respond appropriately at the evaluated decision points. Analysis across behavior-specific benchmarks further reveals weaknesses in how current LLMs behave as agents. By turning deployment problems into targeted benchmarks, TraceDance could serve as a key component of the recursive self-improvement (RSI) loop.