Can LLMs Credibly Transform the Creation of Panel Data from Diverse Historical Tables?

📅 2025-05-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses high parsing error rates (40%), prohibitive costs, and poor reproducibility in digitizing historical county-level motor vehicle registration records. We propose the first multimodal large language model (MLLM)-driven, end-to-end table parsing framework tailored for historical economic data. Our method integrates OCR enhancement, structured multi-stage prompting, and statistical consistency verification to enable fully automated, verifiable conversion from table images to panel data. Experiments demonstrate a reduction in parsing error rate to 0.3%, an R² of 98.6%, and statistically indistinguishable estimates of key economic indicators—including vehicle penetration growth and persistence—compared to human-annotated ground truth. Processing cost is reduced by 99%. The framework significantly enhances accuracy, efficiency, and accessibility in constructing historical economic datasets, thereby facilitating broader empirical research and cross-disciplinary collaboration.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Language Grounding & Multi-modal NLPComputer Vision: Multi-modal Vision

Application Category

Economics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labelingSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Multimodal LLMs offer a watershed change for the digitization of historical tables, enabling low-cost processing centered on domain expertise rather than technical skills. We rigorously validate an LLM-based pipeline on a new panel of historical county-level vehicle registrations. This pipeline is 100 times less expensive than outsourcing, reduces critical parsing errors from 40% to 0.3%, and matches human-validated gold standard data with an $R^2$ of 98.6%. Analyses of growth and persistence in vehicle adoption are statistically indistinguishable whether using LLM or gold standard data. LLM-based digitization unlocks complex historical tables, enabling new economic analyses and broader researcher participation.
Problem

Research questions and friction points this paper is trying to address.

Can LLMs accurately convert historical tables into panel data
Does LLM-based pipeline reduce costs and errors in data processing
Can LLM-digitized data match human-validated gold standard quality
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal LLMs enable low-cost historical table digitization
LLM-based pipeline reduces parsing errors to 0.3%
LLM digitization matches gold standard data with 98.6% R²
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
V
Ver'onica Backer-Peral
Massachusetts Institute of Technology
V
Vitaly Meursault
Federal Reserve Bank of Philadelphia
Christopher Severen
Christopher Severen
Federal Reserve Bank of Philadelphia
Urban EconomicsEnvironmental EconomicsDevelopment Economics