Who Warmed the Archives? LLMs Overestimate Historical Warmth

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the systematic warm bias exhibited by large language models (LLMs) when extracting climate data from historical archives, demonstrating that reliance solely on high correlation metrics for validation can lead to misuse. We evaluate the reliability of six mainstream LLMs in extracting temperature indices from German-language historical documents. By comparing lexical baselines, fine-tuned Transformers, and date-stripping ablation experiments, we identify the root cause of the observed biases. Our findings reveal that all models exhibit a significant warm bias that increases over time. This error stems from modern climatic priors rather than deficiencies in temporal inference capabilities, effectively doubling the error of the best-performing model. Consequently, this work challenges the single-metric validation paradigm and demonstrates that employing LLMs for historical climate reconstruction necessitates the integration of multi-dimensional verification mechanisms to ensure data fidelity.
📝 Abstract
Historical archives are an under-used source for extending the instrumental climate record backward in time, and LLMs offer a way to extract the indices climatologists derive by hand. Beyond measuring how well systems extract this signal, we check whether their errors are safe to use for cross-century comparison, since a good correlation score does not rule out systematic, era-linked bias. Comparing lexical baselines, fine-tuned historical transformers, and LLM prompting on the Pfister temperature index across five centuries of German text, lexical methods beat every fine-tuned transformer we test, including one pretrained from scratch on historical German (r=-0.016). All six LLMs we test (Gemini 2.5 Flash, GPT-5-mini, DeepSeek v4 Flash, Claude Sonnet 4.6, Qwen3.7-Plus, Kimi-K2.6-Fast) show a warm bias that grows with calendar year, with the same sign in every model (slopes +0.13 to +0.34/century, p<0.01). The effect is modest in size (r-squared approx equal to 0.01 to 0.05) but consistent across six independently developed models. The best-correlated of the six, Gemini 2.5 Flash, matches the best lexical correlation (r=0.32) at double the error. An ablation stripping explicit dates and calendar-era markers from the quotes leaves this trend essentially unchanged, favoring an anachronistic present-day prior over the model correctly inferring the quote's era. Correlation alone is thus insufficient for vetting an LLM as a historical-climate-index oracle.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Historical Climate Archives
Systematic Bias
Temperature Index Extraction
Anachronistic Bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
Historical Climatology
Systematic Bias
Ablation Study
Pfister Temperature Index
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Claudiu Creanga
Interdisciplinary School of Doctoral Studies, HLT Research Center, University of Bucharest, Romania
Liviu P. Dinu
Liviu P. Dinu
Professor, University of Bucharest, Dept. of Computer Science,
Computational LinguisticsNatural Language ProcessingComputational Historical Linguistics