Score
Design and implement data-scoring and filtering pipelines that compute character error rate (CER) between reference and predicted/transcribed strings and use those CER scores to identify, remove, weight, or subsample noisy labeled sequence examples. Build thresholding and selection policies and integrate them into pretraining or fine-tuning workflows, and analyze how CER-based filtering changes dataset quality and downstream model performance.
Traditional character error rate (CER) fails to accurately evaluate the overall performance of end-to-end OCR systems when page layout parsing is erroneous. To address this limitation, this work proposes the Character Error Vector (CEV), which for the first time enables decomposability in OCR evaluation by disentangling errors into three components: layout parsing, text recognition, and their interaction. Instantiated within a bag-of-characters framework, the CEV yields two concrete metrics—Spatially Aware Character Error Rate (SpACER) and a Jensen–Shannon divergence–based character distribution measure—effectively bridging the gap between layout analysis and localized OCR evaluation while supporting unified assessment across diverse annotation schemes. Experiments on historical newspaper datasets demonstrate that CEV not only reveals the superiority of conventional pipeline approaches over current end-to-end models but also accurately predicts dominant error sources with an F1 score of 0.91 using only easily obtainable thresholds.
This work addresses the performance limitations of Arabic handwritten text recognition (HTR) caused by pervasive label noise in training data—including transcription, segmentation, orientation, and non-text errors. The authors propose CER-HV, a novel framework that integrates character error rate (CER)-driven sample ranking with human-in-the-loop validation to systematically detect and clean label noise in multilingual Arabic datasets. Leveraging a CRNN-based CER estimator with early stopping, the method identifies erroneous samples with 90% precision on the Muharaf dataset and 80–86% on PHTI. After cleaning with CER-HV, HTR models achieve a 1.0–1.8% absolute CER reduction and establish state-of-the-art results across multiple benchmarks, offering a general and efficient paradigm for denoising high-noise handwritten text data.
Traditional automatic speech recognition evaluation metrics, such as word error rate (WER) and character error rate (CER), fail to capture human perception of errors and neglect linguistic and semantic influences. This work proposes a novel paradigm that embeds any perception-oriented evaluation metric into the minimum edit distance (minED) framework to produce an intuitively interpretable equivalent error rate. For the first time, this approach translates human perceptual modeling into a comprehensible error rate format, enabling quantification of error severity from the perspective of human understanding. The resulting metric not only aligns closely with human judgments but also effectively identifies recognition errors that critically impact semantic comprehension.
This work addresses the challenge of noisy labels and poor domain alignment in large-scale weakly supervised speech recognition data, which often limits model performance. The authors propose a three-stage training strategy: first, pretraining on the full weakly supervised dataset; second, refining the model through a second pretraining phase using a high-quality subset selected based on character error rate (CER); and third, fine-tuning on samples from this subset that are acoustically similar to the target domain, as measured by an acoustic similarity metric. By jointly incorporating data filtering and domain-relevant sample selection without introducing external data, the method significantly improves recognition accuracy. Experiments on 90,000 hours of Japanese weakly supervised data demonstrate CER reductions of up to 6.4% and 4.0% under the proposed approach.
To address the scarcity of labeled data for domain adaptation of automatic speech recognition (ASR) models in resource-constrained scenarios, this paper proposes a multi-stage pseudo-label filtering framework that jointly leverages word error rate (WER) prediction, named entity recognition (NER), and character error rate (CER) analysis to significantly improve pseudo-label quality and selection robustness. The method generates initial pseudo-labels using Whisper (an encoder-decoder architecture) and Zipformer (a transcription-focused model), then filters 100 hours of high-quality utterances—just 1.4% of a 7,500-hour customer-service speech corpus. Fine-tuning an ASR model on this compact subset achieves a WER of 12.3%, matching performance attained when training on the full pseudo-labeled set. This framework establishes a scalable, low-cost paradigm for efficient ASR domain adaptation in low-resource, high-noise settings—particularly suitable for small-scale organizations deploying ASR in domains such as customer service.
This study addresses the performance degradation of pathology report classification models when deployed across different cancer registries due to distributional shift. The authors propose a reproducible, end-to-end supervised classification framework that constructs training and evaluation data closely mirroring real-world deployment scenarios. This is achieved through institution-stratified sampling, separate processing of registry-linked case reports, and blinded estimation of positive rates and label noise. By integrating supervised text classification with threshold optimization constrained by a target false negative rate (FNR), the model achieves substantially improved generalization. Evaluated on a test set of 418,000 reports, the Kentucky model attains an FNR of 0.003, an FPR of 0.097, and an F1 score of 0.922—significantly outperforming baseline approaches, which achieved an F1 score of 0.860.
This study investigates the impact of label noise on the generalization performance of language models, particularly in low-resource or high-noise settings. Focusing on Russian multi-domain text classification, it presents the first systematic comparison between Confident Learning and Dataset Cartography as automated methods for detecting label errors. Leveraging a fine-tuned rubert-base-cased model, the authors apply these techniques to filter noisy training data and validate their efficacy through controlled random-deletion baselines. Results demonstrate that Confident Learning substantially improves macro-F1 scores on small, high-noise datasets, whereas Dataset Cartography adopts a more conservative approach, removing fewer samples. Both methods consistently outperform random deletion, with their relative effectiveness closely dependent on dataset size and noise level.
This work addresses the inconsistency and poor reproducibility of character error rate (CER) and word error rate (WER) metrics in evaluating automatic transcription models, which stem from opaque preprocessing and ambiguous definitions of characters and words in existing text alignment tools. To resolve this, the authors propose Stringalign, a lightweight Python library that introduces FAIR principles to transcription evaluation for the first time. Stringalign enables reproducible, fine-grained error analysis through Unicode-aware transparent normalization, flexible tokenization strategies, and character- and word-level alignment algorithms. Coupled with interactive visualizations, Stringalign significantly enhances evaluation consistency and interpretability across OCR, handwritten text recognition (HTR), and automatic speech recognition (ASR) tasks, thereby facilitating effective model diagnosis and selection.
This work addresses the challenge of scarce alignment data for post-trained open models. We propose a semantic similarity–based data extraction method: leveraging high-quality embedding models to measure semantic distances between model outputs and original supervised fine-tuning (SFT) or reinforcement learning (RL) training data, thereby identifying and reconstructing latent alignment samples. Our approach reveals that knowledge distillation may implicitly reproduce original training data—a previously underappreciated privacy and copyright risk—and achieves, for the first time, scalable, high-fidelity alignment data extraction from black-box post-trained models. Experiments demonstrate that fine-tuning base models solely on extracted data significantly restores performance across critical dimensions—including long-context reasoning, safety, instruction following, and mathematical reasoning—validating the feasibility of effective reverse engineering and reuse of alignment data. This establishes a novel paradigm for model interpretability, data provenance, and safety evaluation.
This work addresses the challenge of efficiently selecting high-quality subsets from ultra-large corpora to reduce fine-tuning costs. The authors propose CRAFT, a method that decomposes the source–target joint distribution and employs a two-stage strategy: first allocating the selection budget across clusters according to k-means proportions to match the source distribution of the validation set, then selecting within each cluster samples whose target embeddings best approximate the validation target distribution. CRAFT is the first approach to combine clustering with conditional expectation distance for data selection. Theoretically, proportional allocation bounds the KL divergence regardless of the specific embedding scheme. Evaluated on fine-tuning mBART with 33 million sentence pairs, CRAFT achieves a BLEU score of 43.34—outperforming TSDS by 2.13 points—and runs 2.8× faster than TAROT, completing the entire pipeline on CPU in under one minute.