Handwritten Text Recognition Lives in the High-Pixel Variance Subspace

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear rationale for selecting self-supervised pretraining methods in handwritten text recognition (HTR) and the observed superiority of pixel reconstruction over contrastive learning. We propose a high-variance subspace theory revealing that discriminative signals concentrate within high-variance pixel subspaces, thereby elucidating the intrinsic connection between pixel reconstruction and feature alignment. Extensive multilingual benchmarking evaluates six self-supervised approaches, including masked image modeling and JEPA, coupled with large language model decoders. Results demonstrate that pixel-oriented methods consistently achieve the lowest character error rates across all benchmarks. Furthermore, frozen encoders yield performance comparable to fully fine-tuned supervised baselines, firmly establishing the superiority of pixel-oriented self-supervised pretraining for HTR.
📝 Abstract
In self-supervised pretraining for Handwritten Text Recognition (HTR), pixel reconstruction methods outperform contrastive methods, unlike in natural-image classification. We argue that this difference follows from where discriminative signal lies in pixel space: for HTR, it is concentrated in high-variance directions and largely absent from low-variance ones. This predicts that objectives preserving high-variance pixel content will transfer best. We test six SSL methods from three families (pixel-grounded MIM, JEPA, and contrastive) under matched encoder, data, and evaluation protocols on six handwriting benchmarks across five languages. With full labels, pixel-groundrounded SSL achieves the lowest CER on every benchmark and both frozen probes, exposes per-position character information that other families recover only through the readout, and is the only family to benefit from pretraining on real handwriting. Pixel-grounded representations are also more label efficient. Across datasets, encoder alignment with the high-variance pixel subspace predicts CER within every method. With a pretrained LLM decoder, a frozen pixel-grounded encoder is competitive with fully fine-tuned supervised baselines; full fine-tuning achieves the lowest mean CER and ranks first or second on every benchmark. These results show that the value of pixel reconstruction depends on where discriminative signal lies in the input.
Problem

Research questions and friction points this paper is trying to address.

Handwritten Text Recognition
Self-Supervised Learning
Pixel Reconstruction
High-Variance Subspace
Representation Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Handwritten Text Recognition
Self-Supervised Learning
High-Variance Subspace
Masked Image Modeling
Pixel Reconstruction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.