🤖 AI Summary
Automatic plant identification in tropical regions—exemplified by the Guiana Shield in South America, hosting ~1,000 highly diverse plant species—is severely constrained by the scarcity of field-collected image data.
Method: This study systematically evaluates, for the first time, the feasibility of leveraging digitized herbarium specimens to enhance cross-domain recognition performance. We formulate a large-scale herbarium-to-field-image classification task and employ deep convolutional neural networks to learn feature mappings between herbarium sheets and in-situ photographs, enabling effective knowledge transfer.
Contribution/Results: By integrating state-of-the-art models from multiple research teams, we demonstrate that combining a small number of field images with abundant herbarium data significantly improves classification accuracy on real-world photographs. This work bridges classical botanical resources with modern deep learning, establishing a scalable, transferable technical paradigm for intelligent biodiversity monitoring in data-scarce regions.
📝 Abstract
Automated identification of plants has improved considerably thanks to the recent progress in deep learning and the availability of training data with more and more photos in the field. However, this profusion of data only concerns a few tens of thousands of species, mostly located in North America and Western Europe, much less in the richest regions in terms of biodiversity such as tropical countries. On the other hand, for several centuries, botanists have collected, catalogued and systematically stored plant specimens in herbaria, particularly in tropical regions, and the recent efforts by the biodiversity informatics community made it possible to put millions of digitized sheets online. The LifeCLEF 2020 Plant Identification challenge (or "PlantCLEF 2020") was designed to evaluate to what extent automated identification on the flora of data deficient regions can be improved by the use of herbarium collections. It is based on a dataset of about 1,000 species mainly focused on the South America's Guiana Shield, an area known to have one of the greatest diversity of plants in the world. The challenge was evaluated as a cross-domain classification task where the training set consist of several hundred thousand herbarium sheets and few thousand of photos to enable learning a mapping between the two domains. The test set was exclusively composed of photos in the field. This paper presents the resources and assessments of the conducted evaluation, summarizes the approaches and systems employed by the participating research groups, and provides an analysis of the main outcomes.