AnthroDial: Benchmarking LLM Anthropomorphism in Autonomous Social Interaction

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of autonomous communication decision-making and unified evaluation frameworks for large language models (LLMs) in open-ended social interactions. To this end, we propose AnthroDial, an integrated framework encompassing interaction, evaluation, and training. Specifically, it introduces a MindFlow dynamic buffering mechanism to enable autonomous asynchronous interaction and establishes CAPS-Eval, a multidimensional evaluation system that quantifies anthropomorphism across cognitive, affective, and behavioral dimensions. Furthermore, the SEEDS environment expansion strategy and the DiAPO adaptive optimization approach are designed to enhance model capabilities. Experimental results demonstrate that the proposed method significantly improves the interaction autonomy and naturalness of LLMs, while validating both the reliability of the evaluation system and the effectiveness of the training paradigm.
📝 Abstract
Large language models (LLMs) are increasingly deployed as social agents, yet credible human-like interaction requires more than fluent responses or persona consistency. Agents must autonomously decide whether, when, and how to communicate while adapting to evolving contexts, goals, and relationships. Existing research, however, lacks a unified approach to enabling, evaluating, and improving such capabilities in continuous, open-ended interaction. We introduce AnthroDial, a unified framework for developing anthropomorphic social agents from three complementary aspects: MindFlow, a lightweight interaction harness that enables autonomous, asynchronous, and adaptive communication through a dynamic Mind Buffer; CAPS-Eval, a theory-grounded framework for evaluating cognitive, affective, and behavioral dimensions of anthropomorphic interaction; and a scalable training paradigm that combines SEEDS for environment expansion with DiAPO for adaptive capability optimization. We further construct evaluation datasets covering everyday communication, game interaction, and long-horizon character interaction. Extensive experiments across diverse models and scenarios demonstrate improved interaction autonomy and naturalness, validate the reliability, discriminativeness, and agreement with human rankings of CAPS-Eval, and confirm the effectiveness of our training paradigm. Together, these components provide a unified framework for developing credible human-like social agents in open-ended interaction.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Anthropomorphism
Social Agents
Autonomous Interaction
Benchmarking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Anthropomorphism
Social Agents
Autonomous Interaction
Evaluation Framework
Large Language Models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
W
Wentao Liu
Shanghai Institute of Innovation
Xi Chen
Xi Chen
Chengdu University of Technology
Micro-nanomotorsActive colloidsMembrane separation
S
Siyu Song
East China Normal University
B
Biao Yuan
Anhui University
Y
Yu Zhang
Shanghai Jiaotong University
Z
Zhou Zhuotong
Fudan University
J
Jingying Zhou
University of Melbourne
G
Guohao Feng
Shanghai Jiaotong University
S
Shasha Hu
Zhejiang University
T
Tianfu Wang
The Hong Kong University of Science and Technology (Guangzhou)
Shangshang Yang
Shangshang Yang
Anhui University
Trustworthy Intelligent EducationAutoMLEvolutionary Computation
H
Haoyang Liu
Shanghai Tianyou Software Co., Ltd.
Y
Youjia Li
Shanghai Tianyou Software Co., Ltd.
Xiaokun Wang
Xiaokun Wang
Nanjing University
Video analytics
Min Ji
Min Ji
Academy of Mathematics and Systems Science, CAS
partial differential equationsgeometric analysisstochastic analysis
J
Ji Wang
Zhejiang Century Huatong Group Co., Ltd.