🤖 AI Summary
This study addresses high parsing error rates (40%), prohibitive costs, and poor reproducibility in digitizing historical county-level motor vehicle registration records. We propose the first multimodal large language model (MLLM)-driven, end-to-end table parsing framework tailored for historical economic data. Our method integrates OCR enhancement, structured multi-stage prompting, and statistical consistency verification to enable fully automated, verifiable conversion from table images to panel data. Experiments demonstrate a reduction in parsing error rate to 0.3%, an R² of 98.6%, and statistically indistinguishable estimates of key economic indicators—including vehicle penetration growth and persistence—compared to human-annotated ground truth. Processing cost is reduced by 99%. The framework significantly enhances accuracy, efficiency, and accessibility in constructing historical economic datasets, thereby facilitating broader empirical research and cross-disciplinary collaboration.
📝 Abstract
Multimodal LLMs offer a watershed change for the digitization of historical tables, enabling low-cost processing centered on domain expertise rather than technical skills. We rigorously validate an LLM-based pipeline on a new panel of historical county-level vehicle registrations. This pipeline is 100 times less expensive than outsourcing, reduces critical parsing errors from 40% to 0.3%, and matches human-validated gold standard data with an $R^2$ of 98.6%. Analyses of growth and persistence in vehicle adoption are statistically indistinguishable whether using LLM or gold standard data. LLM-based digitization unlocks complex historical tables, enabling new economic analyses and broader researcher participation.