🤖 AI Summary
This study addresses the long-standing fragmentation of tropical species data across multiple biodiversity platforms and the lack of systematic linkage between Chinese common names and international trade regulations such as CITES. We present a cross-domain dataset encompassing 410,499 tropical species, integrating three subdomains—plants, aquatic organisms, and pets—and introduce a novel trade life-cycle–oriented ontology. The dataset establishes a traceable layer of Chinese common names and provides direct mappings from each species to CITES Species+ entries. By harmonizing taxonomic information from six authoritative sources—including GBIF, POWO, and iNaturalist—and implementing a four-tier confidence mechanism for nomenclature management, the resource achieves 99.50% coverage of Chinese common names. Released under a CC-BY 4.0 license on Zenodo, this dataset offers a high-integrity foundational resource for scientific research, regulatory enforcement, and public applications.
📝 Abstract
We describe a versioned cross-domain dataset of 410,499 active tropical species (working snapshot 2026-04-20) spanning three applied subdomains -- tropical_plants, tropical_aquatic, and tropical_pets -- that share a commercial and regulatory life cycle but are distributed across kingdom-organised biodiversity infrastructures. The resource joins taxonomic identifiers from GBIF, Plants of the World Online, iNaturalist, NCBI Taxonomy, the Catalogue of Life and the Encyclopedia of Life, and adds three original layers: a cross-domain ontology that re-segments taxa along trade and husbandry contexts; a Chinese vernacular layer with explicit per-name provenance under a typology that excludes unverified machine-generated proposals; and a CITES source-linkage layer connecting each taxon to its Species+ entry. Chinese vernacular coverage -- the proportion of taxa carrying a CJK Chinese name distinct from the scientific binomial -- reaches 99.50 percent (408,456 of 410,499; full-population count). Coverage characterises completeness, not name-translation accuracy; the latter is bounded by the four-level provenance typology and is the subject of a preliminary internal review reported here, with a blind external audit identified as the principal open item. Upstream content is referenced by stable identifier only for the original-contribution layers, supporting CC-BY 4.0 reuse. The dataset is deposited on Zenodo (10.5281/zenodo.20377811). This preprint is the canonical v1.0 description of the dataset's current state; future Data Descriptor submission is anticipated but is contingent on the validation and release-engineering items listed in the Limitations.