ALSO: Adversarial Online Strategy Optimization for Social Agents

πŸ“… Unknown Date
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Existing large language model–driven social agents rely on static role definitions, limiting their ability to dynamically adapt to evolving interlocutors and contexts in non-stationary multi-turn dialogues. This work proposes the first online policy optimization framework tailored for social agents, formulating multi-turn social interaction as an adversarial multi-armed bandit problem. The action space combines fixed role descriptions with dynamic policy instructions, while a lightweight neural proxy predicts sparse rewards from interaction history, enabling continual adaptation without assuming environmental stationarity. Evaluated on the Sotopia benchmark, the proposed method significantly outperforms both static baselines and existing optimization strategies, demonstrating superior adaptability and robustness in complex social interactions.
πŸ“ Abstract
Social simulation provides a compelling testbed for studying social intelligence, where agents interact through multi-turn dialogues under evolving contexts and strategically adapting opponents. Such environments are inherently non-stationary, requiring agents to dynamically adjust their strategies over time. However, most Large Language Model (LLM) based social agents rely on static personas, while existing approaches for enhancing social intelligence, such as offline reinforcement learning or external planners, are ill-suited to these settings, typically assuming stationarity and incurring substantial training overhead. To bridge this gap, we propose \textbf{ALSO} (\textbf{A}dversarial on\textbf{L}ine \textbf{S}trategy \textbf{O}ptimization), the first framework for online strategy optimization in multi-agent social simulation. ALSO advances social adaptation through two key contributions. (1) ALSO formulates multi-turn interaction as an adversarial bandit problem, where combinations of static personas and dynamic strategy instructions are treated as arms, providing a principled solution to non-stationarity without relying on environmental stability assumptions. (2) To predict rewards and generalize sparse feedback in multi-turn dialogues, ALSO introduces a lightweight neural surrogate to predict rewards from interaction histories, enabling sample-efficient exploration and continuous online adaptation. Experiments on the Sotopia benchmark demonstrate that ALSO consistently outperforms static baselines and existing optimization methods in dynamic environments, validating the effectiveness of adversarial online strategy optimization for building robust social agents.
Problem

Research questions and friction points this paper is trying to address.

non-stationary environments
social agents
online strategy optimization
multi-agent social simulation
dynamic adaptation
Innovation

Methods, ideas, or system contributions that make the work stand out.

adversarial bandits
online strategy optimization
social agents
non-stationary environments
neural surrogate
πŸ’Ό Related Jobs
No related jobs found.
X
Xiang Li
School of Artificial Intelligence, Tianjin University, Tianjin, China
Liping Yi
Liping Yi
Tenure-Track Associate Professor, Tianjin University
Federated LearningLLM Multi-Agent
M
Mingze Kong
The Chinese University of Hong Kong, Shenzhen, China
Min Zhang
Min Zhang
East China Normal University
Zhongxiang Dai
Zhongxiang Dai
Assistant Professor, The Chinese University of Hong Kong, Shenzhen
Machine LearningData-Centric AILarge Language ModelsMulti-Armed BanditsBayesian Optimization
Q
QingHua Hu
School of Artificial Intelligence, Tianjin University, Tianjin, China