The Character Error Vector: Decomposable errors for page-level OCR evaluation

📅 2026-04-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Traditional character error rate (CER) fails to accurately evaluate the overall performance of end-to-end OCR systems when page layout parsing is erroneous. To address this limitation, this work proposes the Character Error Vector (CEV), which for the first time enables decomposability in OCR evaluation by disentangling errors into three components: layout parsing, text recognition, and their interaction. Instantiated within a bag-of-characters framework, the CEV yields two concrete metrics—Spatially Aware Character Error Rate (SpACER) and a Jensen–Shannon divergence–based character distribution measure—effectively bridging the gap between layout analysis and localized OCR evaluation while supporting unified assessment across diverse annotation schemes. Experiments on historical newspaper datasets demonstrate that CEV not only reveals the superiority of conventional pipeline approaches over current end-to-end models but also accurately predicts dominant error sources with an F1 score of 0.91 using only easily obtainable thresholds.

Technology Category

Machine Learning: Evaluation and AnalysisNatural Language Processing: Interpretability, Analysis, and Evaluation of NLP ModelsSearch and Optimization: Evaluation and Analysis

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsSystems and Infrastructure for Web, Mobile and WoT: Web performance, measurement, and characterizationSecurity and Privacy: Large-scale security measurements
📝 Abstract
The Character Error Rate (CER) is a key metric for evaluating the quality of Optical Character Recognition (OCR). However, this metric assumes that text has been perfectly parsed, which is often not the case. Under page-parsing errors, CER becomes undefined, limiting its use as a metric and making evaluating page-level OCR challenging, particularly when using data that do not share a labelling schema. We introduce the Character Error Vector (CEV), a bag-of-characters evaluator for OCR. The CEV can be decomposed into parsing and OCR, and interaction error components. This decomposability allows practitioners to focus on the part of the Document Understanding pipeline that will have the greatest impact on overall text extraction quality. The CEV can be implemented using a variety of methods, of which we demonstrate SpACER (Spatially Aware Character Error Rate) and a Character distribution method using the Jensen-Shannon Distance. We validate the CEV's performance against other metrics: first, the relationship with CER; then, parse quality; and finally, as a direct measure of page-level OCR quality. The validation process shows that the CEV is a valuable bridge between parsing metrics and local metrics like CER. We analyse a dataset of archival newspapers made of degraded images with complex layouts and find that state-of-the-art end-to-end models are outperformed by more traditional pipeline approaches. Whilst the CEV requires character-level positioning for optimal triage, thresholding on easily available values can predict the main error source with an F1 of 0.91. We provide the CEV as part of a Python library to support Document understanding research.
Problem

Research questions and friction points this paper is trying to address.

Character Error Rate
page-level OCR evaluation
parsing errors
Document Understanding
OCR quality assessment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Character Error Vector
decomposable evaluation
page-level OCR
SpACER
Jensen-Shannon Distance
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jonathan Bourne
M
Mwiza Simbeye
J
Joseph Nockels