🤖 AI Summary
This study addresses the fragmentation in AI game commentary research and the inability of conventional holistic evaluations to capture functional heterogeneity. To overcome these limitations, this work constructs a unified cross-game benchmark spanning board games, sports, and esports, and proposes a genre-aware structured evaluation framework. Standardized assessment is achieved through multimodal data alignment, a human annotation protocol, and reliability verification algorithms. The research reveals that real-time observation and strategic analysis constitute the primary performance bottlenecks in current systems. Ultimately, this work provides a comparable and interpretable diagnostic foundation for advancing AI commentary systems.
📝 Abstract
Game commentary is an open-ended generation task requiring multimodal perception, strategic reasoning, and contextual knowledge. Existing AI-Generated Game Commentary (AI-GGC) studies remain fragmented across games, modalities, and evaluation protocols, while overlap-based or holistic evaluators fail to capture the functional heterogeneity of commentary. We introduce \textsc{GameCommBench}, a unified benchmark spanning board games, sports, and esports, with commentary aligned to heterogeneous game contexts and annotated by commentary type. We further propose Type-Aware Commentary Evaluation (TACE), a structured framework for evaluating different types of commentary. We then validate TACE for reliability and human agreement, and use it to benchmark representative AI commentators. Results reveal non-uniform capability profiles, with live observation and strategic analysis emerging as major bottlenecks. Together, \textsc{GameCommBench} and TACE provide a diagnostic foundation for comparable and interpretable AI-GGC evaluation.