TaxonRL: Reinforcement Learning with Intermediate Rewards for Interpretable Fine-Grained Visual Reasoning

📅 2026-03-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge faced by conventional vision-language models in fine-grained classification, particularly their difficulty in distinguishing visually similar species within the same genus or family. To overcome this limitation, the authors propose a reinforcement learning framework based on Group Relative Policy Optimization that structures the classification process as a hierarchical reasoning pipeline—progressing from species to genus to family—and incorporates an intermediate reward mechanism to guide the model toward generating interpretable and verifiable decision paths. By uniquely integrating hierarchical classification with intermediate rewards, the method achieves a state-of-the-art average accuracy of 91.7% on the Birds-to-Words dataset, substantially outperforming human experts (77.3%) and demonstrating strong cross-domain generalization capabilities on primate and marine species classification tasks.

Technology Category

Computer Vision: Language and VisionNatural Language Processing: Language Grounding & Multi-modal NLPMachine Learning: Multi-class/Multi-label Learning & Extreme Classification

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Traditional vision-language models struggle with contrastive fine-grained taxonomic reasoning, particularly when distinguishing between visually similar species within the same genus or family. We introduce TaxonRL, a reinforcement learning approach using Group Relative Policy Optimization with intermediate rewards that decomposes the reasoning process into hierarchical taxonomic predictions. Our method incentivizes models to explicitly reason about species-level, genus-level, and family-level features before making final classifications. This structured approach is designed not only to boost accuracy but also to yield a transparent, verifiable decision-making process. On the challenging Birds-to-Words dataset, TaxonRL achieves 91.7\% average accuracy, exceeding human performance (77.3\%) while generating interpretable reasoning traces. We demonstrate strong cross-domain generalization, showing substantial gains in primate and marine species verification. Our results establish that enforcing structured, hierarchical reasoning provides a powerful and transferable framework for fine-grained visual discrimination.
Problem

Research questions and friction points this paper is trying to address.

fine-grained visual reasoning
taxonomic reasoning
visually similar species
contrastive classification
interpretability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Intermediate Rewards
Hierarchical Reasoning
Fine-Grained Visual Classification
Interpretable AI
M
Maximilian von Klinski
Hasso Plattner Institute, University of Potsdam, Germany
M
Maximilian Schall
Hasso Plattner Institute, University of Potsdam, Germany