AgriScope: Pixel-Grounded Multimodal Understanding for Agricultural Images

📅 2026-09-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出AgriScope框架,通过生物语义编码和密集空间表示解决农业图像中植物疾病、害虫识别等问题,实现像素级理解。
📝 Abstract
Agricultural image understanding requires fine-grained recognition of plant diseases, pests, crop structures, and botanical species under complex real-world conditions. Despite recent advances in Multimodal Large Language Models (MLLMs), existing models remain limited to text-only outputs and lack pixel-level visual grounding capabilities. In this work, we introduce AgriScope, a unified pixel-grounded multimodal framework for agricultural image understanding. AgriScope jointly supports image-level, region-level, and pixel-level understanding within a unified framework, enabling tasks such as grounded caption generation, referring expression segmentation, and multi-turn multimodal interaction for agricultural imagery. AgriScope integrates biologically specialized semantic representations with dense spatial grounding through biological-semantic encoding, dense spatial representations, and pixel decoding. To support large-scale grounded learning, we introduce AgriGround, a large-scale pixel-grounded agricultural multimodal instruction-tuning dataset containing over 500K images and 11M instruction-following samples spanning plant disease analysis, crop and weed identification, insect pest recognition, and fine-grained botanical understanding. AgriGround is constructed through a multi-stage automatic annotation pipeline that integrates multimodal caption generation, phrase-level grounding, segmentation mask generation, and task-oriented instruction synthesis to produce densely grounded supervision. Extensive experiments across multiple agricultural vision-language tasks demonstrate the effectiveness of AgriScope in pixel-grounded multimodal understanding, establishing a strong benchmark for agricultural vision-language learning and visual grounding. The dataset and code will be made publicly available at (https://github.com/boudiafA/AgriScope)
Problem

Research questions and friction points this paper is trying to address.

Agricultural Image Understanding
Pixel-Level Visual Grounding
Multimodal Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

pixel-grounded multimodal understanding
agricultural image understanding
biological-semantic encoding
dense spatial representations
AgriGround dataset
🔎 Similar Papers
No similar papers found.