🤖 AI Summary
This work addresses the lack of effective and scalable evaluation methodologies for foundational models in social robotics. It proposes a three-tiered evaluation funnel paradigm that integrates general-purpose metrics, simulated interactions, and real-world robot experiments. For the first time, it defines five core evaluation dimensions for social robot foundation models and establishes a cross-tier community-driven collaborative evaluation framework. The approach systematically elucidates the mapping relationships between these dimensions across the three evaluation tiers, identifies critical gaps in current assessment practices, and fosters the co-development of standardized evaluation protocols. By doing so, it provides both theoretical grounding and practical pathways to guide the selection and optimization of foundational models in social robotics.
📝 Abstract
Researchers who seek to build social robot applications on foundation models are faced with a difficult question: how should we pick a model? Public leaderboards offer little guidance: the demands of real-time, embodied social interaction lie largely outside their focus. And direct evaluation is impractical at scale: each embodied study requires scarce participant, robot, and experimenter time. In this paper, we identify five evaluation dimensions for foundation models in social robots: (i) conversational competence, (ii) user safety, (iii) embodied character, (iv) target scene effectiveness, and (v) audience appropriateness. To make model selection cheaper and better informed, we propose a three-tiered evaluation funnel paradigm that first filters with general metrics, then extends to simulated interactions, and terminates in more expensive, robot-specific evaluation. We map all five dimensions across all three tiers, chart where applicable evaluation methods exist and are missing, and close with a call to action: let's build the evaluation framework together as a community.