🤖 AI Summary
This study challenges the validity of using self-reported personality inventories (e.g., BFI, HEXACO) to characterize the intended personality designs of LLM-based chatbots. We systematically evaluated 500 distinct personality-configured LLM agents, assessing the alignment between their self-assessed scores and human-perceived personality traits—as well as objective interaction quality metrics (e.g., task completion rate, naturalness, trustworthiness). Results reveal negligible correlations (r < 0.2) between self-reports and both user perceptions and interaction outcomes, indicating severe deficiencies in criterion and predictive validity. Crucially, task context and interaction modality emerged as significant moderating factors. Consequently, we propose— for the first time—a novel “contextualized, interaction-driven” paradigm for personality assessment in LLM agents. This paradigm rejects static self-reporting in favor of a dynamic, behavior-based validation framework grounded in authentic conversational interactions.
📝 Abstract
Personality design plays an important role in chatbot development. From rule-based chatbots to LLM-based chatbots, evaluating the effectiveness of personality design has become more challenging due to the increasingly open-ended interactions. A recent popular approach uses self-report questionnaires to assess LLM-based chatbots' personality traits. However, such an approach has raised serious validity concerns: chatbot's"self-report"personality may not align with human perception based on their interaction. Can LLM-based chatbots"self-report"their personality? We created 500 chatbots with distinct personality designs and evaluated the validity of self-reported personality scales in LLM-based chatbot's personality evaluation. Our findings indicate that the chatbot's answers on human personality scales exhibit weak correlations with both user perception and interaction quality, which raises both criterion and predictive validity concerns of such a method. Further analysis revealed the role of task context and interaction in the chatbot's personality design assessment. We discuss the design implications for building contextualized and interactive evaluation of the chatbot's personality design.