cer-based data filtering

Design and implement data-scoring and filtering pipelines that compute character error rate (CER) between reference and predicted/transcribed strings and use those CER scores to identify, remove, weight, or subsample noisy labeled sequence examples. Build thresholding and selection policies and integrate them into pretraining or fine-tuning workflows, and analyze how CER-based filtering changes dataset quality and downstream model performance.

cer-baseddatafiltering

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.22
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Traditional character error rate (CER) fails to accurately evaluate the overall performance of end-to-end OCR systems when page layout parsing is erroneous. To address this limitation, this work proposes the Character Error Vector (CEV), which for the first time enables decomposability in OCR evaluation by disentangling errors into three components: layout parsing, text recognition, and their interaction. Instantiated within a bag-of-characters framework, the CEV yields two concrete metrics—Spatially Aware Character Error Rate (SpACER) and a Jensen–Shannon divergence–based character distribution measure—effectively bridging the gap between layout analysis and localized OCR evaluation while supporting unified assessment across diverse annotation schemes. Experiments on historical newspaper datasets demonstrate that CEV not only reveals the superiority of conventional pipeline approaches over current end-to-end models but also accurately predicts dominant error sources with an F1 score of 0.91 using only easily obtainable thresholds.

Character Error RateDocument UnderstandingOCR quality assessment

This work addresses the performance limitations of Arabic handwritten text recognition (HTR) caused by pervasive label noise in training data—including transcription, segmentation, orientation, and non-text errors. The authors propose CER-HV, a novel framework that integrates character error rate (CER)-driven sample ranking with human-in-the-loop validation to systematically detect and clean label noise in multilingual Arabic datasets. Leveraging a CRNN-based CER estimator with early stopping, the method identifies erroneous samples with 90% precision on the Muharaf dataset and 80–86% on PHTI. After cleaning with CER-HV, HTR models achieve a 1.0–1.8% absolute CER reduction and establish state-of-the-art results across multiple benchmarks, offering a general and efficient paradigm for denoising high-noise handwritten text data.

Arabic-script HTRCharacter Error Ratedata quality

Traditional automatic speech recognition evaluation metrics, such as word error rate (WER) and character error rate (CER), fail to capture human perception of errors and neglect linguistic and semantic influences. This work proposes a novel paradigm that embeds any perception-oriented evaluation metric into the minimum edit distance (minED) framework to produce an intuitively interpretable equivalent error rate. For the first time, this approach translates human perceptual modeling into a comprehensible error rate format, enabling quantification of error severity from the perspective of human understanding. The resulting metric not only aligns closely with human judgments but also effectively identifies recognition errors that critically impact semantic comprehension.

Automatic Speech RecognitionCharacter Error Rateevaluation metrics

This work addresses the challenge of noisy labels and poor domain alignment in large-scale weakly supervised speech recognition data, which often limits model performance. The authors propose a three-stage training strategy: first, pretraining on the full weakly supervised dataset; second, refining the model through a second pretraining phase using a high-quality subset selected based on character error rate (CER); and third, fine-tuning on samples from this subset that are acoustically similar to the target domain, as measured by an acoustic similarity metric. By jointly incorporating data filtering and domain-relevant sample selection without introducing external data, the method significantly improves recognition accuracy. Experiments on 90,000 hours of Japanese weakly supervised data demonstrate CER reductions of up to 6.4% and 4.0% under the proposed approach.

automatic speech recognitiondomain specificitylarge-scale datasets

Efficient Data Selection for Domain Adaptation of ASR Using Pseudo-Labels and Multi-Stage Filtering

Jun 04, 2025
PR
Pradeep Rangappa
🏛️ Idiap Research Institute | Uniphore Systems | EPFL | University of Zurich | Brno University

To address the scarcity of labeled data for domain adaptation of automatic speech recognition (ASR) models in resource-constrained scenarios, this paper proposes a multi-stage pseudo-label filtering framework that jointly leverages word error rate (WER) prediction, named entity recognition (NER), and character error rate (CER) analysis to significantly improve pseudo-label quality and selection robustness. The method generates initial pseudo-labels using Whisper (an encoder-decoder architecture) and Zipformer (a transcription-focused model), then filters 100 hours of high-quality utterances—just 1.4% of a 7,500-hour customer-service speech corpus. Fine-tuning an ASR model on this compact subset achieves a WER of 12.3%, matching performance attained when training on the full pseudo-labeled set. This framework establishes a scalable, low-cost paradigm for efficient ASR domain adaptation in low-resource, high-noise settings—particularly suitable for small-scale organizations deploying ASR in domains such as customer service.

Filtering pseudo-labels for high-quality training segmentsImproving ASR adaptation with limited labeled dataReducing dataset size while maintaining performance

Latest Papers

What's happening recently
View more

This study addresses the performance degradation of pathology report classification models when deployed across different cancer registries due to distributional shift. The authors propose a reproducible, end-to-end supervised classification framework that constructs training and evaluation data closely mirroring real-world deployment scenarios. This is achieved through institution-stratified sampling, separate processing of registry-linked case reports, and blinded estimation of positive rates and label noise. By integrating supervised text classification with threshold optimization constrained by a target false negative rate (FNR), the model achieves substantially improved generalization. Evaluated on a test set of 418,000 reports, the Kentucky model attains an FNR of 0.003, an FPR of 0.097, and an F1 score of 0.922—significantly outperforming baseline approaches, which achieved an F1 score of 0.860.

biomedical NLPcancer registrylabel noise

This study investigates the impact of label noise on the generalization performance of language models, particularly in low-resource or high-noise settings. Focusing on Russian multi-domain text classification, it presents the first systematic comparison between Confident Learning and Dataset Cartography as automated methods for detecting label errors. Leveraging a fine-tuned rubert-base-cased model, the authors apply these techniques to filter noisy training data and validate their efficacy through controlled random-deletion baselines. Results demonstrate that Confident Learning substantially improves macro-F1 scores on small, high-noise datasets, whereas Dataset Cartography adopts a more conservative approach, removing fewer samples. Both methods consistently outperform random deletion, with their relative effectiveness closely dependent on dataset size and noise level.

data qualitylabel errorslanguage models

This work addresses the inconsistency and poor reproducibility of character error rate (CER) and word error rate (WER) metrics in evaluating automatic transcription models, which stem from opaque preprocessing and ambiguous definitions of characters and words in existing text alignment tools. To resolve this, the authors propose Stringalign, a lightweight Python library that introduces FAIR principles to transcription evaluation for the first time. Stringalign enables reproducible, fine-grained error analysis through Unicode-aware transparent normalization, flexible tokenization strategies, and character- and word-level alignment algorithms. Coupled with interactive visualizations, Stringalign significantly enhances evaluation consistency and interpretability across OCR, handwritten text recognition (HTR), and automatic speech recognition (ASR) tasks, thereby facilitating effective model diagnosis and selection.

automatic transcription evaluationcharacter error rateevaluation transparency

Extracting alignment data in open models

Oct 21, 2025
FB
Federico Barbero
🏛️ University of Oxford | National University of Singapore | OpenAI | Google DeepMind | Anthropic | MentaLeap | AI Sequrity Company

This work addresses the challenge of scarce alignment data for post-trained open models. We propose a semantic similarity–based data extraction method: leveraging high-quality embedding models to measure semantic distances between model outputs and original supervised fine-tuning (SFT) or reinforcement learning (RL) training data, thereby identifying and reconstructing latent alignment samples. Our approach reveals that knowledge distillation may implicitly reproduce original training data—a previously underappreciated privacy and copyright risk—and achieves, for the first time, scalable, high-fidelity alignment data extraction from black-box post-trained models. Experiments demonstrate that fine-tuning base models solely on extracted data significantly restores performance across critical dimensions—including long-context reasoning, safety, instruction following, and mathematical reasoning—validating the feasibility of effective reverse engineering and reuse of alignment data. This establishes a novel paradigm for model interpretability, data provenance, and safety evaluation.

Extracting alignment training data from post-trained modelsInvestigating risks of data regurgitation in distillation practicesUsing embedding models to identify semantic data similarities

This work addresses the challenge of efficiently selecting high-quality subsets from ultra-large corpora to reduce fine-tuning costs. The authors propose CRAFT, a method that decomposes the source–target joint distribution and employs a two-stage strategy: first allocating the selection budget across clusters according to k-means proportions to match the source distribution of the validation set, then selecting within each cluster samples whose target embeddings best approximate the validation target distribution. CRAFT is the first approach to combine clustering with conditional expectation distance for data selection. Theoretically, proportional allocation bounds the KL divergence regardless of the specific embedding scheme. Evaluated on fine-tuning mBART with 33 million sentence pairs, CRAFT achieves a BLEU score of 43.34—outperforming TSDS by 2.13 points—and runs 2.8× faster than TAROT, completing the entire pipeline on CPU in under one minute.

data selectionfine-tuninglarge-scale corpora

Hot Scholars

DB

Diego Belzarena

Universidad de la República, Uruguay
Computer VisionElectrical Engineering
DM

Danilo Mandic

Prof. of Machine Intelligence, Dept of Electrical and Electronic Eng., Imperial College London, UK
Machine Intelligence and Statistical Signal Proc.Biomedicine and FinanceHearables and Ear-EEGDeep RNNs
YK

Yohei Kawaguchi

Hitachi, Ltd.
Acoustic Signal ProcessingSignal ProcessingMachine LearningSpeech Processing
AN

Anguelos Nicolaou

University of Graz
Computer VisionNeural NetworksDocument Image AnalysisDigital Humanities
MG

Marina Gardella

Centre Borelli, ENS Paris-Saclay, Université Paris-Saclay
Image processing