From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models

πŸ“… 2026-07-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Existing benchmarks exhibit limitations in document granularity, task design, and data provenance, hindering comprehensive evaluation of large language models (LLMs) in multi-granularity event analysis. To address this gap, this work proposes MiGUE-Benchβ€”the first unified evaluation benchmark spanning single-document to cross-document multi-granularity event analysis, encompassing four core tasks: event detection, relation reasoning, structural induction, and future prediction. A novel LLM-driven self-correcting annotation framework, MiGUE-Pipeline, integrated with retrieval-augmented generation (RAG), enables high-quality, large-scale automatic construction of event-centric datasets. Systematic evaluation of mainstream LLMs on this benchmark reveals their performance boundaries and critical weaknesses in complex event understanding, thereby offering clear directions for future model improvement.
πŸ“ Abstract
Event analysis is an essential and fundamental direction of information extraction, involving various event-centric tasks at different granularity of documents. While large language models (LLMs) have preliminarily achieved promising performance in part of these tasks individually, their capability in event analysis still lacks comprehensive understanding due to restricted document granularity, task designs, and data source of existing benchmarks. To address these limitations, we introduce MiGUE-Bench, a systematic benchmark for assessing the performance of LLMs in multi-granularity event analysis. To support large-scale evaluation, we first develop an LLM-driven self-correcting annotation framework called MiGUE-Pipeline, enabling scalable acquisition of high-quality source data of events with automatic labels. Then, we design four core tasks in our benchmark, i.e., event detection, relation reasoning, structure induction, and future prediction, to probe model competence at different levels, from atomic event details to complex cross-document narratives. Extensive experiments on state-of-the-art LLMs and retrieval-augmented generation (RAG) methods delineate the current capability boundary and identify critical deficiencies, providing insights into the future improvement of LLMs in challenging event analysis tasks.
Problem

Research questions and friction points this paper is trying to address.

event analysis
large language models
multi-granularity
benchmark
information extraction
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-granularity event analysis
LLM benchmarking
self-correcting annotation
cross-document reasoning
event structure induction
πŸ”Ž Similar Papers
No similar papers found.
Tao Wen
Tao Wen
The University of Manchester
NetworksDecision makingData ScienceNonlinear dynamicsGame theory
Shuai Shao
Shuai Shao
University of Science and Technology of China
Complexity TheoryInformation Theory
Pei Ke
Pei Ke
Associate Professor, University of Electronic Science and Technology of China
Natural Language ProcessingNatural Language GenerationDialogue SystemLarge Language Model
Xu Han
Xu Han
Research Assistant Professor, Tsinghua University
Natural Language ProcessingLarge Language ModelKnowledge GraphInformation Extraction
Jie Zou
Jie Zou
University of Electronic Science and Technology of China
Information RetrievalNatural Language ProcessingRecommender SystemsMultimedia
G
Guannan Li
Laboratory of Intelligent Collaborative Computing, University of Electronic Science and Technology of China, Chengdu, China
T
Tao Tian
Laboratory of Intelligent Collaborative Computing, University of Electronic Science and Technology of China, Chengdu, China
J
Jinjie Qiu
Laboratory of Intelligent Collaborative Computing, University of Electronic Science and Technology of China, Chengdu, China
Lan Wang
Lan Wang
Professor of Computer Science, University of Memphis
computer networks
K
Ke Qin
Laboratory of Intelligent Collaborative Computing, University of Electronic Science and Technology of China, Chengdu, China