Exploring In-Context Learning for Handwritten Text Recognition

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the domain shift challenges inherent in historical document digitization, where traditional handwritten text recognition (HTR) relies heavily on extensive annotated data and model fine-tuning. We construct a transcription pipeline leveraging the in-context learning capabilities of pretrained vision-language models (VLMs), providing the first empirical validation that general-purpose VLMs can achieve zero-shot HTR through in-context learning. Furthermore, this work reveals the trade-off between context size and error rate in cross-domain scenarios. Experimental results demonstrate that, without any parameter updates, the proposed approach exhibits competitive performance comparable to conventional HTR models across both in-domain and cross-domain settings.
📝 Abstract
Handwritten Text Recognition (HTR) systems have become an indispensable tool for the digitization of historical documents. Not only do they cut down time and cost, but they also allow democratizing access and processing of their contents by generating their transcripts. However, literature in HTR currently focuses mostly on specialized models that require large amounts of annotated samples to achieve satisfactory performance. We explore the use of In-Context Learning with pre-trained Vision-Language Models (VLMs) to create a transcription pipeline without updating the model's parameters. We then evaluate this pipeline across multiple collections and models, and demonstrate that general-purpose VLMs can be effectively taught how to transcribe handwritten text from images. To assess how our observations may translate to practical applications, we evaluate the performance in a Cross-Domain (CD) scenario, where context examples are drawn from a different collection than the query image. Results in both the controlled In-Domain (ID) scenario and the realistic CD scenario follow the same patterns. First, as context size grows, the error range is expected to narrow towards the average performance. Thus, larger context sizes sacrifice the performance of the oracle-best sampling for lower expected error rates. The results obtained show that, without any parameter updates, this methodology has strong potential to compete with traditional HTR in the presence of domain shift. Moreover, we show and argue that some context samplings work better than others and suggest more effort should be put into finding an ideal sampling method in future work.
Problem

Research questions and friction points this paper is trying to address.

Handwritten Text Recognition
In-Context Learning
Vision-Language Models
Cross-Domain
Domain Shift
Innovation

Methods, ideas, or system contributions that make the work stand out.

In-Context Learning
Vision-Language Models
Handwritten Text Recognition
Parameter-free
Cross-Domain
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.