Score
Design and build models, feature‑learning methods, and end‑to‑end pipelines that assign species‑level labels among visually similar categories by learning subtle discriminative features; this includes techniques to aggregate segmented patches or region proposals into object‑level species assignments and to produce high‑precision species predictions for individual objects.
Existing animal classification models (e.g., SpeciesNet) typically yield coarse-grained taxonomic labels (e.g., order or class), limiting species-level identification. To address this, we propose a five-stage hierarchical reclassification framework that integrates EfficientNetV2-M and CLIP-based visual embeddings, augmented with centroid clustering, triplet-loss-driven metric learning, and an adaptive cosine-distance scoring mechanism—enabling fine-grained mapping from high-level taxa to species. Innovatively, we introduce a bird-priority coverage strategy and a high-confidence filtering mechanism to enhance discriminative robustness. Evaluated on the LILA BC Desert Lion Conservation dataset, our method recovers 761 bird detections and refines 456 ambiguous coarse labels, achieving an overall accuracy of 96.5%, with 64.9% of predictions precisely resolved at the species level—substantially improving both accuracy and practical utility for wildlife image recognition.
To address cross-granularity prediction errors in hierarchical image classification caused by visual inconsistency at test time, this paper proposes the first hierarchical classification paradigm grounded in *intra-image visual consistency*. Our method requires no external semantic supervision or pixel-level annotations; instead, it employs a self-supervised segmentation alignment mechanism to visually align fine-grained predictions with coarse-grained regions within the same image. By integrating multi-scale feature modeling with CLIP’s zero-shot transfer capability, the framework enforces both semantic and visual consistency. Evaluated on multiple hierarchical classification benchmarks, our approach significantly outperforms zero-shot CLIP and existing state-of-the-art methods, achieving higher classification accuracy and improved prediction coherence. Notably, it simultaneously enhances unsupervised image segmentation quality—thereby strengthening model interpretability and robustness—without additional supervision.
This work addresses the challenge of identity confusion in wild multi-species animal re-identification, which arises from variations in pose, illumination, background, resolution, and morphology. To tackle this issue, the authors propose a species-aware graph construction framework that integrates foreground-aware preprocessing, species-specific backbone networks, LightGlue for local feature matching, and LightGBM for pairwise scoring. Robust clustering is achieved through conservative edge insertion followed by Leiden community detection, effectively mitigating over-merging caused by bridging edges. The method was evaluated in the AnimalCLEF 2026 competition, where it ranked 5th among 230 teams, achieving an Adjusted Rand Index (ARI) of 0.733 on the public test set and 0.674 on the private test set, demonstrating its effectiveness and state-of-the-art performance.
This work addresses the limited robustness of multimodal species classification in large-scale wild datasets, which often stems from neglecting the hierarchical structure inherent in biological taxonomy. To this end, the authors propose an end-to-end hierarchical-aware multimodal learning framework that explicitly models taxonomic hierarchies. The approach introduces Hierarchical Information Regularization (HiR) to refine the geometric structure of the embedding space and incorporates a lightweight fusion predictor capable of both unimodal and joint inference. Evaluated on multiple large-scale biodiversity benchmarks, the method outperforms strong multimodal baselines by over 14% in accuracy, demonstrating particularly superior performance under challenging conditions such as missing modalities or degraded DNA barcodes.
To address the cross-modal identification challenge of insect species—including unknown taxa—in large-scale biodiversity monitoring, this paper proposes the first multimodal framework integrating images, DNA barcodes, and taxonomic text. Methodologically, it innovatively adapts CLIP-style contrastive learning for image–DNA cross-modal alignment, enabling zero-shot species recognition without task-specific fine-tuning. The framework jointly encodes visual features via ResNet/ViT, models DNA sequences using k-mer representations and Transformers, and embeds taxonomic labels, all unified within a shared multimodal embedding space to achieve semantic alignment across the three modalities. Evaluated on real-world field data, the approach achieves a zero-shot classification accuracy 8.3% higher than the best unimodal baseline, significantly improving generalization to both known and novel insect species. This work establishes a new paradigm for automated, scalable, and dynamic biodiversity monitoring.
This work addresses the challenge faced by conventional vision-language models in fine-grained classification, particularly their difficulty in distinguishing visually similar species within the same genus or family. To overcome this limitation, the authors propose a reinforcement learning framework based on Group Relative Policy Optimization that structures the classification process as a hierarchical reasoning pipeline—progressing from species to genus to family—and incorporates an intermediate reward mechanism to guide the model toward generating interpretable and verifiable decision paths. By uniquely integrating hierarchical classification with intermediate rewards, the method achieves a state-of-the-art average accuracy of 91.7% on the Birds-to-Words dataset, substantially outperforming human experts (77.3%) and demonstrating strong cross-domain generalization capabilities on primate and marine species classification tasks.
This work addresses the challenges of automatic marine species classification in underwater imagery, particularly cross-platform domain shift, high visual similarity among closely related species, and inconsistent annotation granularity. To tackle these issues, the authors propose a deep learning framework that explicitly incorporates biological taxonomic hierarchy during both training and inference. The approach innovatively integrates taxonomy-aware weighted loss, minimum-risk Bayesian inference, multi-scale feature encoding, and a disentangled hierarchical classification head. This design enhances the model’s generalization under distributional shifts and improves recognition performance across multiple taxonomic granularities. Evaluated on the FathomNet 2025 dataset, the method achieves a mean classification distance of 1.581, approaching the current state-of-the-art result of 1.535, thereby demonstrating its effectiveness and robustness.
研究测试了2-8B参数范围内的视觉-语言模型在边缘设备上识别物种的能力,与专业模型BioCLIP对比,发现所有模型在野外图像上的表现均下降,表明数据训练的重要性。
This work addresses two key challenges in biodiversity research: the fine-grained identification of visually similar species and the automatic discovery of unknown species in open environments. To this end, the authors propose DeepTaxon, a novel framework that formulates species discovery as an explicit retrieval-based decision problem for the first time. By integrating retrieval-augmented multimodal contrastive reasoning with chain-of-thought and interpretable comparative mechanisms, DeepTaxon unifies identification and discovery into a single evidence-driven decision process grounded in retrieved references. Notably, it generates supervisory signals automatically without human annotation, enabling joint optimization of both tasks. Experiments demonstrate that DeepTaxon significantly outperforms existing methods on a large-scale in-distribution benchmark and six out-of-distribution datasets, exhibiting exceptional zero-shot transfer capability, test-time scalability, and cross-encoder stability.
This study addresses the high cost and low efficiency of manual annotation in ecological image analysis by proposing a label-free, species-level automatic clustering method. The authors establish the first systematic benchmark framework to evaluate zero-shot clustering performance across 60 species using five Vision Transformer models—including DINOv2 and DINOv3—combined with various dimensionality reduction and clustering algorithms. Experimental results demonstrate that the proposed approach achieves a V-measure of 0.958 (0.943 under fully unsupervised settings) at the species level, requiring expert review of only 1.14% of images. Moreover, the method effectively uncovers ecologically meaningful intraspecific structures, such as variations in sex, age, and coat color. To support broader adoption, the authors release an open-source toolkit enabling ecologists to efficiently analyze large-scale wildlife image datasets.