From Pixels to Structure: Lightweight Vision-Language Models for Document OCR and Structured JSON Extraction

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the data privacy concerns, high costs, and energy consumption associated with deploying commercial closed-source vision-language models (VLMs) for heritage archive digitization. We propose an autonomous and controllable structured extraction framework based on a lightweight open-source VLM. Methodologically, utilizing a 7B-parameter base model, we design a constraint-aware protocol that integrates classical image preprocessing with multi-stage fine-tuning techniques, while quantifying the independent contribution of each module to document OCR-to-JSON extraction accuracy. Experimental results demonstrate that the adapted lightweight model can efficiently replace manual annotation or black-box systems, achieving compliant and precise structured data extraction. This work presents a novel paradigm that balances performance with controllability for the digitization of sensitive archival materials.
📝 Abstract
While massive, closed-source Vision-Language Models (VLMs) set strong benchmarks for document understanding, their dependence on commercial APIs limits adoption in institutional archives due to data autonomy concerns, recurring costs, and the environmental footprint of hyperscale computing. This is especially acute in heritage digitization, where documents include historical handwriting, domain-specific terminology (e.g., jewelry, prehistory, architecture), and non-standard layouts requiring high-dimensional structured extraction. We present a comparative study of eight open-source lightweight VLMs (up to 7B parameters) for Optical Character Recognition (OCR)-to-structure across three university heritage collections. Given a document image, models must extract text and generate schema-compliant JSON, enabling automatic validation and downstream use. We evaluate models under a constraint-aware protocol across zero-shot, few-shot, and fine-tuning settings, measuring extraction fidelity and structured-output quality using Character Error Rate (CER), Approximate Normalized Levenshtein Similarity (ANLS*), and mean Average Precision F1 (mAP-F1). Against a fine-tuning baseline, we further test the independent impact of (i) hyperparameter optimization, (ii) classical image preprocessing (illumination flattening, denoising, and CLAHE), and (iii) multi-stage training. Finally, we analyze the trade-off between dataset-specific fine-tuning and a single multi-dataset checkpoint, where joint training enables one model to operate across collections but can shift performance between datasets. Overall, we show that carefully adapted VLMs with up to 7B parameters can provide a sustainable, private, high-performing alternative to manual transcription or commercial black-box systems, and we offer actionable guidance for heritage institutions seeking institution-controlled OCR-to-JSON extraction.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Document OCR
Structured JSON Extraction
Heritage Digitization
Lightweight Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
Structured JSON Extraction
Lightweight OCR
Document Understanding
Heritage Digitization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.