🤖 AI Summary
This study addresses the lack of autonomous communication decision-making and unified evaluation frameworks for large language models (LLMs) in open-ended social interactions. To this end, we propose AnthroDial, an integrated framework encompassing interaction, evaluation, and training. Specifically, it introduces a MindFlow dynamic buffering mechanism to enable autonomous asynchronous interaction and establishes CAPS-Eval, a multidimensional evaluation system that quantifies anthropomorphism across cognitive, affective, and behavioral dimensions. Furthermore, the SEEDS environment expansion strategy and the DiAPO adaptive optimization approach are designed to enhance model capabilities. Experimental results demonstrate that the proposed method significantly improves the interaction autonomy and naturalness of LLMs, while validating both the reliability of the evaluation system and the effectiveness of the training paradigm.
📝 Abstract
Large language models (LLMs) are increasingly deployed as social agents, yet credible human-like interaction requires more than fluent responses or persona consistency. Agents must autonomously decide whether, when, and how to communicate while adapting to evolving contexts, goals, and relationships. Existing research, however, lacks a unified approach to enabling, evaluating, and improving such capabilities in continuous, open-ended interaction. We introduce AnthroDial, a unified framework for developing anthropomorphic social agents from three complementary aspects: MindFlow, a lightweight interaction harness that enables autonomous, asynchronous, and adaptive communication through a dynamic Mind Buffer; CAPS-Eval, a theory-grounded framework for evaluating cognitive, affective, and behavioral dimensions of anthropomorphic interaction; and a scalable training paradigm that combines SEEDS for environment expansion with DiAPO for adaptive capability optimization. We further construct evaluation datasets covering everyday communication, game interaction, and long-horizon character interaction. Extensive experiments across diverse models and scenarios demonstrate improved interaction autonomy and naturalness, validate the reliability, discriminativeness, and agreement with human rankings of CAPS-Eval, and confirm the effectiveness of our training paradigm. Together, these components provide a unified framework for developing credible human-like social agents in open-ended interaction.