🤖 AI Summary
This study addresses the systematic warm bias exhibited by large language models (LLMs) when extracting climate data from historical archives, demonstrating that reliance solely on high correlation metrics for validation can lead to misuse. We evaluate the reliability of six mainstream LLMs in extracting temperature indices from German-language historical documents. By comparing lexical baselines, fine-tuned Transformers, and date-stripping ablation experiments, we identify the root cause of the observed biases. Our findings reveal that all models exhibit a significant warm bias that increases over time. This error stems from modern climatic priors rather than deficiencies in temporal inference capabilities, effectively doubling the error of the best-performing model. Consequently, this work challenges the single-metric validation paradigm and demonstrates that employing LLMs for historical climate reconstruction necessitates the integration of multi-dimensional verification mechanisms to ensure data fidelity.
📝 Abstract
Historical archives are an under-used source for extending the instrumental climate record backward in time, and LLMs offer a way to extract the indices climatologists derive by hand. Beyond measuring how well systems extract this signal, we check whether their errors are safe to use for cross-century comparison, since a good correlation score does not rule out systematic, era-linked bias. Comparing lexical baselines, fine-tuned historical transformers, and LLM prompting on the Pfister temperature index across five centuries of German text, lexical methods beat every fine-tuned transformer we test, including one pretrained from scratch on historical German (r=-0.016). All six LLMs we test (Gemini 2.5 Flash, GPT-5-mini, DeepSeek v4 Flash, Claude Sonnet 4.6, Qwen3.7-Plus, Kimi-K2.6-Fast) show a warm bias that grows with calendar year, with the same sign in every model (slopes +0.13 to +0.34/century, p<0.01). The effect is modest in size (r-squared approx equal to 0.01 to 0.05) but consistent across six independently developed models. The best-correlated of the six, Gemini 2.5 Flash, matches the best lexical correlation (r=0.32) at double the error. An ablation stripping explicit dates and calendar-era markers from the quotes leaves this trend essentially unchanged, favoring an anachronistic present-day prior over the model correctly inferring the quote's era. Correlation alone is thus insufficient for vetting an LLM as a historical-climate-index oracle.