HistoriQA-ThirdRepublic: Multi-Hop Question Answering Corpus for Historical Research, Parliamentary Debates from the French Third Republic (1870-1940)

📅 2026-06-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of natural language processing benchmarks that support the integration of cross-source, temporal, and sparse evidence—a critical need in historical research—by introducing a French multi-hop question answering corpus grounded in parliamentary debates and press coverage from the French Third Republic. Comprising 1,782 questions, this dataset is the first to incorporate direct collaboration with historians, explicitly modeling authentic multi-hop reasoning patterns observed in real-world historical inquiry. Through rigorous source selection and alignment, question validation, and metadata integration, the project delivers a domain-specific evaluation resource tailored for retrieval-augmented and large language models. Furthermore, it offers a methodological framework readily adaptable to archival materials in other languages and national contexts, thereby effectively bridging the gap between NLP capabilities and the practical demands of historical scholarship.
📝 Abstract
We present HistoriQA-ThirdRepublic: a French-language dataset of multi-hop historical questions derived from parliamentary debates and newspapers of the French Third Republic. Designed in collaboration with a historian, the corpus captures complex reasoning patterns typical of historical inquiry, including cross-source synthesis, temporal reasoning, and the integration of sparse evidence. The dataset is made of 1782 questions and emphasizes multi-hop connections across heterogeneous historical documents, providing a resource for evaluating retrieval-augmented and large language model systems in domain-specific contexts. We describe the methodology for constructing the corpus, including the selection and alignment of sources, question validation, and metadata integration. While the dataset focuses on French historical documents, our methodology can be readily adapted to other languages and national corpora. Finally, we demonstrate how the corpus can support realistic evaluation scenarios for multi-hop question answering, bridging the gap between NLP benchmarks and the needs of historical scholarship.
Problem

Research questions and friction points this paper is trying to address.

multi-hop question answering
historical research
parliamentary debates
heterogeneous documents
temporal reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-hop question answering
historical corpus
retrieval-augmented models
temporal reasoning
cross-source synthesis