🤖 AI Summary
This study addresses the reliance on handcrafted rules and poor cross-task transferability of existing user behavior simulators by proposing the SWORD framework. Without requiring domain priors or task-specific engineering, SWORD jointly optimizes multi-agent workflow topologies and natural language prompts using only scalar metrics. Furthermore, it leverages text gradients to enable unsupervised feature importance discovery and autonomous rule mining. Experimental results demonstrate that SWORD significantly outperforms baseline methods while utilizing less data and smaller models. By achieving superior predictive accuracy at minimal API cost and automatically identifying domain-relevant signals, this work presents an efficient new paradigm for general-purpose behavior simulation.
📝 Abstract
User behavior simulation is the computational modeling of user interactions within information systems through the use of simulated agents in place of live users. It supports system testing and evaluation, decision-making and forecasting, and user experience design. Existing simulators rely on hand-crafted rules or domain expertise that transfers poorly across tasks. SWORD (Simulation-driven Workflow and Prompt Optimization with Role-based Design) is introduced as a framework that jointly optimizes multi-agent workflow topology and natural-language prompts. It is guided solely by a scalar task metric, without domain initialization or task-specific engineering. The experimental results demonstrate that SWORD achieves statistically significant gains over prompt-only, workflow-only, and staged-optimization baselines under a controlled, identical-backbone comparison. Against the strongest published domain-specific baseline, SWORD further improves accuracy while using a smaller backbone model, substantially less training data, and a very reasonable API cost (\$4--\$6 for each dataset). Beyond predictive performance, SWORD autonomously discovers domain-relevant signals, review-sentiment mapping rules and epidemiological decay priors, purely from scalar error feedback, establishing textual gradients as a mechanism for unsupervised feature-importance discovery in user behavior modeling.