conduct archival analysis

Designs and implements workflows to locate, retrieve, sample, digitize, and curate collections of archival documents and organizational records, producing cleaned, annotated document collections and associated metadata. Builds analyses and syntheses from those collections—annotating structure and fields, normalizing formats and layouts, constructing timelines, triangulating documentary evidence, and packaging datasets with provenance and sampling strategies for reuse.

conductarchivalanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$124K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing retrieval-augmented generation (RAG) systems struggle to provide holistic interpretations of historical digital collections and lack the ability to dynamically integrate expert knowledge for complex queries. This work proposes a conversational document analysis system that, for the first time, combines dynamic knowledge graphs with conversational RAG. During user interactions, the system incrementally constructs a graph structure that fuses archival expert knowledge with retrieved results, serving as external memory for the language model. By moving beyond the traditional RAG paradigm—which relies solely on raw documents—this approach effectively models cross-document relationships, long-range dependencies, and implicit knowledge, substantially enhancing the system’s capacity to answer multi-record complex questions and deliver deeper historical interpretations.

Conversational RetrievalDocument Collection InterpretationExpert Knowledge Integration

Bridging the Digital Divide: Approach to Documenting Early Computing Artifacts Using Established Standards for Cross-Collection Knowledge Integration Ontology

Jan 22, 2025
MG
Maciej Grzeszczuk
🏛️ Polish-Japanese Academy of Information Technology | The Foundation for the History of Home Computers | University of Maria Curie-Skłodowska

Early computational approaches to digital cultural heritage face fragmentation, semantic inconsistency, and scarcity of professional archival resources. Method: This study pioneers the systematic adoption of the CIDOC Conceptual Reference Model (CIDOC-CRM) ontology in community-driven documentation, proposing a lightweight, extensible ontology construction strategy grounded in participatory design and semantic modeling to enable non-specialist archivists to collaboratively create standardized metadata. Contribution/Results: The project validates CIDOC-CRM’s logical rigor and operational readability in digital archaeology contexts, yielding a reusable minimal viable ontology module set. It has supported over 20 community-based archival initiatives, significantly enhancing cross-institutional knowledge integration and long-term preservation efficacy. The approach establishes a scalable, decentralized governance methodology for digital cultural heritage, offering a transferable paradigm for inclusive, sustainable heritage documentation.

Archiving ChallengesElectronic Device DocumentationHeritage Management

Traditional digitization of historical documents has largely been confined to character transcription, lacking the structural and semantic information necessary for in-depth analysis. This work proposes VERITAS, a framework that reconceptualizes digitization as an integrated pipeline combining transcription, layout analysis, and semantic enrichment. Through four stages—preprocessing, extraction, refinement, and enhancement—it enables end-to-end transformation from document images to structured knowledge. VERITAS employs a model-agnostic, modular design with a schema-driven architecture, allowing declarative specification of information extraction targets and integrating OCR, layout analysis, semantic annotation, and retrieval-augmented generation techniques. Evaluated on over 1,600 pages of Renaissance chronicles, the framework reduces word error rate by 67.6% compared to commercial OCR systems, cuts manual proofreading time by two-thirds, and effectively supports downstream historical question-answering tasks.

computational analysishistorical document digitisationOCR limitations

This study addresses a critical limitation in existing document layout analysis methods, which treat figures and tables as generic objects and thus fail to identify semantically valuable, reusable analytical visual content—referred to as “data snapshots”—in institutional documents. The work introduces the novel task of data snapshot extraction, presents a benchmark dataset comprising humanitarian reports and World Bank policy papers, and proposes an evaluation framework that integrates spatial localization with semantic annotation. Systematic evaluation of multiple open-source layout models reveals consistent shortcomings in handling institutional documents, including confusion between analytical and non-analytical content, fragmentation of composite charts, and lack of contextual awareness. By exposing the generalization bottlenecks of current models in operational documents, this research provides a foundation for future advancements through the public release of its dataset and codebase.

data snapshot extractiondocument layout analysisinstitutional documents

Existing dataset documentation tools struggle to achieve real-world adoption due to ambiguous value propositions, misalignment with practical contexts, insufficient attention to human labor costs, and a lack of systemic integration. This study addresses these challenges through a mixed-methods systematic scoping review of 59 relevant publications, combining qualitative coding with quantitative synthesis to uncover the underlying motivations driving tool design and their relationship to institutional norms. The analysis identifies four key patterns that hinder adoption and advances a responsible AI design perspective that shifts emphasis from individual accountability to institutional solutions. The work advocates embedding sustainable documentation practices within organizational workflows and cultures, offering the HCI community actionable pathways toward institutionalizing responsible data stewardship.

dataset documentationdocumentation practicesResponsible AI

Latest Papers

What's happening recently
View more

Existing Model Cards and Data Cards describe only static models and datasets, lacking documentation of the execution context surrounding generation, transformation, and evaluation processes—thereby limiting reproducibility and bias analysis. This work proposes Workflow Cards, which extend the structured documentation paradigm to dynamic workflow executions for the first time. Built upon provenance data, Workflow Cards generate machine-readable, structured summaries interpretable by both humans and large language models (LLMs), and incorporate a template designed to answer typical execution-related questions. Experimental results demonstrate that Workflow Cards substantially enhance understanding of workflow executions compared to schema-based query interfaces, nearly doubling answer quality and achieving superior performance under both LLM-as-a-Judge and human evaluations.

Data CardsModel Cardsprovenance data

This study addresses the longstanding reliance on manual labor in cataloging digital collections—a process hindered by low efficiency and high costs. The authors systematically evaluate the performance of various artificial intelligence models in automated cataloging tasks, employing both quantitative metrics and qualitative analysis to comprehensively assess accuracy, robustness, and applicability. Their investigation identifies the model architectures best suited for cataloging scenarios and distills a set of transferable, cross-domain principles for AI-driven cataloging. These findings offer both theoretical grounding and practical guidance for cultural heritage institutions seeking to advance their digital transformation through intelligent technologies.

AI modelscataloguingdigital collections

This study addresses the limitations of conventional archival digitization practices, which typically prioritize image scanning and dissemination while neglecting the intrinsic hierarchical structure and archival bond—particularly when handling audiovisual and other multimedia materials, where preserving original organizational logic and semantic relationships proves challenging. To overcome this, the paper introduces the International Image Interoperability Framework (IIIF) into archival science for the first time, integrating archival principles with semantic modeling techniques to develop a digital representation model that respects provenance and original order. This model enables structured, semantically rich expression of hierarchical relationships and archival bonds. The approach is validated through a case study of the “PCI-Unitelefilm” fonds from the AAMOD Foundation, demonstrating its effectiveness in maintaining archival integrity and enhancing semantic interoperability.

archival bondarchivesdigital representation

Hot Scholars

KK

Kevin Klyman

Stanford, Harvard
Foundation ModelsAI RegulationGeopolitics
CT

Christoph Treude

Associate Professor of Computer Science, Singapore Management University
Software EngineeringEmpirical Software EngineeringHuman-AI InteractionAI for Science
RB

Rishi Bommasani

CS PhD, Stanford University
Societal Impact of AIAI PolicyAI GovernanceFoundation Models
JP

Juan Pablo Alperin

Associate Professor, Simon Fraser University
Scholarly CommunicationLatin AmericaOpen AccessOpen Science
LH

Lei Hou

RMIT University
Building Information Modeling (BIM) - Project Management - Construction IT - Productivity Research - Lean Construction