ToolRACER: A Robust Agentic Conversation Emulation Resource for Agent Training and Evaluation

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insufficient robustness of task-oriented dialogue agents in non-cooperative and adversarial scenarios, compounded by a scarcity of high-quality training data. To this end, we propose Sysn, a synthetic data generation framework that leverages multi-agent collaboration to simulate multi-turn interactions among users, assistants, and tools, while injecting adversarial behaviors to produce verified, high-quality data. We construct a synthetic dataset comprising 5.6K trajectories across six domains and introduce the first adversarial multi-turn dialogue benchmark, in which 66% of scenarios are failure-prone, thereby filling a critical gap in evaluating non-cooperative interactions. Experimental results demonstrate that our approach significantly enhances the end-to-end accuracy and robustness of smaller models on function-calling benchmarks such as τ²-bench.
📝 Abstract
Task-oriented conversational agents remain fragile under real world conversation scenarios as they rarely follow a predictable script, especially when users exhibit non-cooperative behavior. Existing function-calling benchmarks often emphasize successful, cooperative interactions and underrepresent adversarial conversation trajectories, thereby limiting the training resources available for developing robust agents. We present ToolRACER, a synthetic data generation pipeline that coordinates user, assistant and tool emulation models to generate and validated multi-turn interactions between a user and an agent. Using \sysn, we construct ToolRACERBench a robust multi-turn conversation benchmark spanning six domains, ranging over 55 varied personas, generating a validated corpus of 5.6K conversation trajectories, with approximately 66\% of conversations containing failure-prone conversation scenarios. We inject adversarial behaviors, producing validated conversational interaction trajectories that capture realistic, robust scenarios. We evaluate models trained on ToolRACERBench against internal benchmarks, as well as on function calling benchmarks such as $τ^2$-bench, BFCLv3 and ACEBench to evaluate agentic accuracy and robustness. Models trained on ToolRACERBench improve end to end agentic accuracy across $τ^2$-bench and ACEBench, demonstrating significant gains when mixed with in-domain dataset in small language models for agent capability tasks.
Problem

Research questions and friction points this paper is trying to address.

task-oriented conversational agents
adversarial conversations
function-calling benchmarks
agent robustness
non-cooperative behavior
Innovation

Methods, ideas, or system contributions that make the work stand out.

Synthetic Data Generation
Adversarial Conversations
Multi-turn Dialogue Benchmark
Function Calling
Agentic Robustness