Quantification of Biodiversity from Historical Survey Text with LLM-based Best-Worst Scaling

📅 2025-02-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Automated quantification of species occurrence frequencies from historical biological survey texts remains challenging due to the poor robustness and high annotation cost of conventional fine-grained multiclass classification approaches for numerical estimation. Method: This work introduces, for the first time, a Best-Worst Scaling (BWS) framework reformulated as an LLM-driven regression task—transforming discrete frequency judgments into relative comparative learning. Contribution/Results: Evaluated on real historical texts using DeepSeek-V3, GPT-4, and Ministral-8B, the method achieves strong agreement with human annotations (Spearman’s ρ > 0.92 for DeepSeek-V3 and GPT-4), outperforming multiclass baselines by 18.7% reduction in MAE. It delivers high estimation accuracy while substantially reducing annotation dependency, establishing a scalable, LLM-powered paradigm for digitizing historical ecological data.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: SummarizationData Mining & Knowledge Management: Conversational Systems for Recommendation & Retrieval

Application Category

Economics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labelingSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
In this study, we evaluate methods to determine the frequency of species via quantity estimation from historical survey text. To that end, we formulate classification tasks and finally show that this problem can be adequately framed as a regression task using Best-Worst Scaling (BWS) with Large Language Models (LLMs). We test Ministral-8B, DeepSeek-V3, and GPT-4, finding that the latter two have reasonable agreement with humans and each other. We conclude that this approach is more cost-effective and similarly robust compared to a fine-grained multi-class approach, allowing automated quantity estimation across species.
Problem

Research questions and friction points this paper is trying to address.

Species frequency estimation from historical texts
Best-Worst Scaling with Large Language Models
Cost-effective automated biodiversity quantification
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-based Best-Worst Scaling
Species frequency quantification
Cost-effective automated estimation
💼 Related Jobs
No related jobs found.
T
Thomas Haider
Chair of Computational Humanities, University of Passau
T
Tobias Perschl
Chair of Computational Humanities, University of Passau
Malte Rehbein
Malte Rehbein
Professor of Computational Humanities, Univerisität Passau
Digital HistoryHistorical EcologyEnvironmental HumanitiesCultural Heritage Digitization