Decoding the Past: Explainable Machine Learning Models for Dating Historical Texts

📅 2025-11-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the precise temporal dating of historical texts in cultural heritage, proposing an interpretable tree-based framework for chronological classification of English historical documents spanning five centuries. Methodologically, it integrates five complementary feature categories—compression ratio, lexical structure, readability, neologism emergence, and temporal distance—and employs SHAP for fine-grained attributional interpretation. The primary contribution is the empirical identification of the 19th century as a pivotal inflection point in English language evolution. Experimental results demonstrate strong predictive performance: 76.7% accuracy for century-level classification; Top-1/Top-2/Top-10 decade-level accuracies of 26.1%, 96.0%, and 85.8%, respectively; and a maximum AUC-ROC of 94.8%. The model significantly outperforms random baselines, exhibits controlled prediction error, and delivers both robust chronological inference and actionable insights for historical linguistics.

Technology Category

Natural Language Processing: Interpretability, Analysis, and Evaluation of NLP ModelsMachine Learning: Ensemble MethodsPlanning, Routing, and Scheduling: Planning with Language Models

Application Category

Web Mining and Content Analysis: Models for Web evolutionSearch and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchSemantics and Knowledge: Methods, algorithms and applications for the development of semantic models, knowledge graphs and other forms of structured data models with machine-interpretable semantics
📝 Abstract
Accurately dating historical texts is essential for organizing and interpreting cultural heritage collections. This article addresses temporal text classification using interpretable, feature-engineered tree-based machine learning models. We integrate five feature categories - compression-based, lexical structure, readability, neologism detection, and distance features - to predict the temporal origin of English texts spanning five centuries. Comparative analysis shows that these feature domains provide complementary temporal signals, with combined models outperforming any individual feature set. On a large-scale corpus, we achieve 76.7% accuracy for century-scale prediction and 26.1% for decade-scale classification, substantially above random baselines (20% and 2.3%). Under relaxed temporal precision, performance increases to 96.0% top-2 accuracy for centuries and 85.8% top-10 accuracy for decades. The final model exhibits strong ranking capabilities with AUCROC up to 94.8% and AUPRC up to 83.3%, and maintains controlled errors with mean absolute deviations of 27 years and 30 years, respectively. For authentication-style tasks, binary models around key thresholds (e.g., 1850-1900) reach 85-98% accuracy. Feature importance analysis identifies distance features and lexical structure as most informative, with compression-based features providing complementary signals. SHAP explainability reveals systematic linguistic evolution patterns, with the 19th century emerging as a pivot point across feature domains. Cross-dataset evaluation on Project Gutenberg highlights domain adaptation challenges, with accuracy dropping by 26.4 percentage points, yet the computational efficiency and interpretability of tree-based models still offer a scalable, explainable alternative to neural architectures.
Problem

Research questions and friction points this paper is trying to address.

Predicting historical text dates using interpretable machine learning models
Integrating multiple linguistic features for temporal classification accuracy
Evaluating model performance across centuries and decades for text dating
Innovation

Methods, ideas, or system contributions that make the work stand out.

Feature-engineered tree models for temporal text classification
Combining five complementary linguistic feature categories
Explainable ML with SHAP analysis revealing linguistic evolution
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Paulo J. N. Pinto
IEETA - Institute of Electronics and Informatics Engineering of Aveiro, and LASI - Intelligent Systems Associate Laboratory, and DETI - Department of Electronics, Telecommunications and Informatics, University of Aveiro, 3810-193 Aveiro, Portugal
Armando J. Pinho
Armando J. Pinho
IEETA/DETI, University of Aveiro
Data CompressionImage CompressionKolmogorov ComplexityAlgorithmic Information Theory
D
Diogo Pratas
IEETA - Institute of Electronics and Informatics Engineering of Aveiro, and LASI - Intelligent Systems Associate Laboratory, and DETI - Department of Electronics, Telecommunications and Informatics, University of Aveiro, 3810-193 Aveiro, Portugal, and DoV - Department of Virology, University of Helsinki, 00014 Helsinki, Finland