🤖 AI Summary
This study addresses a critical yet overlooked issue in historical document digitization: while existing OCR systems prioritize low character error rates (CER), they often neglect semantic hallucinations—such as altered named entities and contextual substitutions—that compromise content fidelity. Focusing on Uruguayan dictatorship-era microfilm documents, this work systematically reveals that vision-language models (VLMs), despite outperforming traditional OCR in CER and word error rate (WER), frequently introduce semantic distortions including orthographic normalization, fabricated content, and substitution of key entities. Through a combination of quantitative metrics and qualitative human analysis, the research demonstrates that conventional evaluation paradigms fail to capture these subtle but significant inaccuracies, thereby advocating for a new semantic trustworthiness assessment framework that transcends character-level accuracy.
📝 Abstract
Optical Character Recognition (OCR) is a key component in the digitization of historical archives. Recently, Vision-Language Models (VLMs) have emerged as strong alternatives to traditional OCR systems, achieving state-of-the-art performance on standard benchmarks. However, their suitability for archival transcription remains insufficiently understood. In this work, we benchmark traditional OCR systems and VLM-based approaches on the Berrutti dataset, a challenging collection of Uruguayan dictatorship-era documents derived from microfilm scans. While VLMs consistently outperform traditional methods in terms of Character Error Rate (CER) and Word Error Rate (WER), we show that these improvements hide a more complex picture. Through a detailed qualitative analysis, we uncover systematic failure modes that are invisible to standard metrics, including orthographic normalization, spurious content generation, and semantic substitutions that preserve fluency while altering meaning. Errors affecting named entities are particularly critical, as they can introduce substantial semantic distortions with minimal impact on CER and WER. These findings reveal a critical gap between quantitative OCR performance and transcription fidelity in real-world archival settings, and highlight the need for evaluation frameworks that go beyond character-level accuracy to capture the semantic reliability of generated transcriptions.