๐ค AI Summary
This work addresses the lack of effective automated evaluation methods for long-form, in-depth research reports generated by large language models (LLMs) by establishing the Deep Research Evaluation subtask. Grounded in the NTCIR evaluation framework, it introduces the first automated benchmark specifically designed for such reports, integrating dataset construction, multidimensional evaluation metrics, and humanโmachine alignment analysis to systematically assess the performance of participating teams' automated quality evaluation approaches. The initiative aggregated 91 system runs from ten teams, validating the effectiveness of diverse methods through comparisons between automated scores and human-annotated ground truth. By publicly releasing the evaluation data, this project significantly extends the application boundaries of automated LLM evaluation.
๐ Abstract
In this paper, we provide an overview of the NTCIR-19 Automatic Evaluation of LLMs 2 (AEOLLM-2) task. Building on the success of the NTCIR-18 core task AEOLLM, we proposed AEOLLM-2 for NTCIR-19 to further investigate automatic evaluation methods for Large Language Models (LLMs), particularly in long-form text generation scenarios. In AEOLLM-2, we introduced a new subtask, Deep Research Evaluation, which focuses on the automatic evaluation of long-form deep research reports generated by LLMs. Participants developed evaluation methods to automatically assess the quality of these reports, and the performance of each method was measured by comparing its scores against human-annotated ground-truth labels. This year, we received 91 runs from 10 teams in total. This paper describes the background of the task, the dataset construction, the evaluation measures, the participants' methods, and the final evaluation results.