Score
Locating, collecting, and critically interpreting historical primary and secondary sources to reconstruct sequences of events and institutional change and to analyze long-term patterns (e.g., NATO innovation or policy reform outcomes).
This study addresses the lack of systematic research on how historians actually employ visualization as data-driven evidence—a gap that hinders the design of interdisciplinary visualization tools. Through a mixed-methods approach, the authors construct a corpus of 14,021 images, apply semi-automated annotation to 4,831 visualization instances, and integrate expert interviews with boundary object analysis using HiFigAtlas. Drawing on a large-scale dataset of 4,142 historical journal articles, they propose the first hierarchical taxonomy of visualizations in historical scholarship. This taxonomy identifies five distinct roles of visualizations and reveals their usage patterns across subfields and temporal dimensions, along with associated cognitive and practical barriers. The findings provide both an empirical foundation and a theoretical framework for developing visualization tools better aligned with historians’ disciplinary needs.
This study addresses the scarcity of annotated historical Italian news corpora by proposing an unsupervised method to automatically identify major socio-political turning points. Building on a diachronic corpus of approximately 600,000 articles from *La Repubblica* (1985–2000), the approach integrates natural language processing, word embeddings, semantic change modeling, complex network analysis, and tools from statistical physics. For the first time, complex systems theory is applied to diachronic media analysis in the Italian context. Without relying on any prior labels, the method successfully detects abrupt shifts in media discourse corresponding to pivotal events such as the transition from Italy’s First to Second Republic, the Gulf War, and the Kosovo War. This work offers a novel paradigm for digital humanities and computational social science by demonstrating how unlabeled textual data can reveal historically significant societal transformations.
This study addresses the challenges of historical document retrieval—such as linguistic evolution, terminological shifts, and cultural biases—that contribute to inequitable access to digital archives. To bridge this gap, the work integrates information retrieval with cultural analysis by constructing the first inclusive evaluation benchmark for 19th-century English texts, based on the British Library’s BL19 collection. Leveraging expert-crafted queries, paragraph-level relevance annotations, and collaboration with large language models, the project introduces a cross-genre knowledge transfer mechanism that adapts the narrative understanding and semantic richness found in fiction to enhance retrieval in non-fiction scholarly documents. This approach significantly improves retrieval accuracy and establishes a historically aware evaluation paradigm for information retrieval that prioritizes interpretability, transparency, and cultural inclusivity, thereby advancing more equitable and emancipatory knowledge infrastructures.
This study addresses the challenge of systematically extracting implicit narrative signals—such as polarization and misinformation—from public discourse for empirical political narrative analysis. We propose a graph-structured modeling approach grounded in Abstract Meaning Representation (AMR) to automatically identify three core narrative elements: actors, events, and perspective-taking. Integrating narratological situation theory, we design heuristic filtering and network-guided retrieval mechanisms, enabling unsupervised and interpretable political narrative discovery. Crucially, we formalize narrative intelligence as retrievable graph signals for the first time, supporting synergistic distant reading and close reading in narrative reconstruction. Experiments on EU State of the Union addresses (2010–2023) demonstrate the method’s effectiveness in uncovering diachronic ideological narrative evolution paths, with strong validity, reproducibility, and real-world applicability.
Historical research lacks AI tools tailored to its needs, making it difficult to efficiently extract structured data from raw document images. To address this gap, this work introduces the first module of the Chronos system, pioneering an agent-driven, interactive, and customizable paradigm for information extraction that overcomes the limitations of conventional vision-language models (VLMs) with rigid, fixed pipelines. By integrating VLMs with natural language–based interactive agents, the approach enables historians to flexibly design, evaluate, and iteratively refine information extraction workflows for heterogeneous historical documents through intuitive natural language commands. The core module has been open-sourced, empowering researchers to maintain full control over their AI-assisted datafication processes.
Standard retrieval-augmented generation (RAG) architectures struggle to align with the methodological demands of interpretive disciplines such as history, particularly in their handling of evidentiary use and interpretive norms. This work proposes a novel RAG framework tailored for historical research that decouples retrieval from generation, incorporates temporal window constraints, and employs dual-path recall combining keyword-based and semantic search. It further introduces “Zwischentexte”—interpretive intermediate texts—as a novel mechanism to embed historiographical epistemological principles directly into system design. A transparent, contestable relevance assessment is achieved through an LLM-as-judge approach. Evaluated on a dataset of 102,189 articles from Der Spiegel, the framework demonstrably mitigates period-specific terminological bias and weakly relevant retrievals, substantially enhancing the historiographical validity and explanatory power of generated outputs.
This study addresses the absence of natural language processing benchmarks that support the integration of cross-source, temporal, and sparse evidence—a critical need in historical research—by introducing a French multi-hop question answering corpus grounded in parliamentary debates and press coverage from the French Third Republic. Comprising 1,782 questions, this dataset is the first to incorporate direct collaboration with historians, explicitly modeling authentic multi-hop reasoning patterns observed in real-world historical inquiry. Through rigorous source selection and alignment, question validation, and metadata integration, the project delivers a domain-specific evaluation resource tailored for retrieval-augmented and large language models. Furthermore, it offers a methodological framework readily adaptable to archival materials in other languages and national contexts, thereby effectively bridging the gap between NLP capabilities and the practical demands of historical scholarship.
Existing retrieval-augmented generation (RAG) systems struggle to provide holistic interpretations of historical digital collections and lack the ability to dynamically integrate expert knowledge for complex queries. This work proposes a conversational document analysis system that, for the first time, combines dynamic knowledge graphs with conversational RAG. During user interactions, the system incrementally constructs a graph structure that fuses archival expert knowledge with retrieved results, serving as external memory for the language model. By moving beyond the traditional RAG paradigm—which relies solely on raw documents—this approach effectively models cross-document relationships, long-range dependencies, and implicit knowledge, substantially enhancing the system’s capacity to answer multi-record complex questions and deliver deeper historical interpretations.
This study interrogates the prevailing scholarly assumption that “structured” and “free-form” interview styles in Holocaust survivor oral histories constitute a strict binary. Drawing on a corpus of over 1,600 archival testimonies, the research employs discourse segmentation, topic modeling, and large language models to quantitatively assess interview structure across dimensions such as thematic coherence, question-answer dynamics, and question typology. Moving beyond conventional dichotomous frameworks, the findings reveal substantial overlap between the two styles at the level of individual narratives, thereby challenging the notion of mutually exclusive categorization. The work not only establishes a scalable and reproducible computational paradigm for oral history analysis but also opens new avenues for digital humanities scholarship and the design of citizen science platforms.