PB-GRPO: Learning Socially Adaptive LLM Agents from Persona-Driven Simulation with Preference-Batched GRPO

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited social adaptability of large language models (LLMs) in inferring implicit user preferences, compounded by the scarcity of authentic multi-turn interaction data. To this end, it constructs a personality-driven social simulation environment and proposes a preference-batched Group Relative Policy Optimization (GRPO) algorithm. By normalizing advantage functions within buckets of users sharing similar preferences, the method stabilizes training across heterogeneous groups. Furthermore, it integrates an LLM-based user simulator with a satisfaction scoring system to optimize agent policies via dialogue-level feedback. Experimental results demonstrate that the proposed approach significantly outperforms strong baselines within the simulated environment, effectively enhancing both the social behavioral performance and personalized adaptability of LLMs.
📝 Abstract
Building LLMs that behave well socially, not merely correctly, requires Building LLMs that behave well socially, not merely correctly, requires more than producing locally helpful responses. A socially competent agent must infer users'unstated goals, respect their preferences, and adapt as the conversation unfolds. These behaviors are inherently multi-turn and social, making them hard to optimize: real interaction data is scarce, and user preferences are typically latent rather than directly observable. To address these challenges, we build on a persona-driven social simulation environment (consisting of a persona library, LLM-based user simulators, and a user-satisfaction scoring system ranging from [0, 1]), to introduce preference-batched GRPO (PB-GRPO), a post-training algorithm that learns socially adaptive policies from conversation-level feedback. Compared to vanilla GRPO, PB-GRPO computes advantages using a normalization estimated across a bucket of users with similar preferences, stabilizing training across a diverse social population. Empirical evidence shows that PB-GRPO improves models'social behavior over strong reinforcement learning baselines in our simulated environment.
Problem

Research questions and friction points this paper is trying to address.

socially adaptive LLM agents
multi-turn social behavior
latent user preferences
interaction data scarcity
Innovation

Methods, ideas, or system contributions that make the work stand out.

PB-GRPO
Preference-Batched GRPO
Socially Adaptive LLM Agents
Persona-Driven Simulation
Reinforcement Learning
💼 Related Jobs
No related jobs found.
J
Jingquan Wang
Amazon
J
Jun Yin
Amazon
X
Xu Han
Amazon
Yongsheng Mei
Yongsheng Mei
Amazon
Jie Hao
Jie Hao
Amazon
LLMNatural Language ProcessingMachine TranslationStatistics
B
Bin Guo
Amazon