AgriGPT-VL: Agricultural Vision-Language Understanding Suite

📅 2025-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the lack of domain-specific multimodal large language models (MLLMs), high-quality vision–language corpora, and rigorous evaluation protocols in agriculture, this paper introduces Agri-VLM—the first multimodal understanding framework tailored for agricultural applications. Methodologically, it proposes a progressive training paradigm integrating multi-agent synthetic data generation, text grounding, shallow/deep cross-modal alignment, and GRPO-based reinforcement learning fine-tuning. Key contributions include: (1) Agri-3M-VL—the largest publicly available agricultural image–text dataset to date; (2) the AgriBench-VL-4K benchmark for vision-language tasks (e.g., visual question answering) and AgriBench-13K for pure-text evaluation, jointly establishing a dual-track, rigorously designed assessment suite balancing openness and fidelity. Experiments demonstrate that Agri-VLM significantly outperforms general-purpose MLLMs on AgriBench-VL-4K while preserving strong linguistic capabilities on AgriBench-13K, validating both the effectiveness and generalizability of domain-specialized multimodal modeling in agriculture.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Computer Vision: Multi-modal VisionNatural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Search and Retrieval-Augmented AI: Agentic searchSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
Despite rapid advances in multimodal large language models, agricultural applications remain constrained by the scarcity of domain-tailored models, curated vision-language corpora, and rigorous evaluation. To address these challenges, we present the AgriGPT-VL Suite, a unified multimodal framework for agriculture. Our contributions are threefold. First, we introduce Agri-3M-VL, the largest vision-language corpus for agriculture to our knowledge, curated by a scalable multi-agent data generator; it comprises 1M image-caption pairs, 2M image-grounded VQA pairs, 50K expert-level VQA instances, and 15K GRPO reinforcement learning samples. Second, we develop AgriGPT-VL, an agriculture-specialized vision-language model trained via a progressive curriculum of textual grounding, multimodal shallow/deep alignment, and GRPO refinement. This method achieves strong multimodal reasoning while preserving text-only capability. Third, we establish AgriBench-VL-4K, a compact yet challenging evaluation suite with open-ended and image-grounded questions, paired with multi-metric evaluation and an LLM-as-a-judge framework. Experiments show that AgriGPT-VL outperforms leading general-purpose VLMs on AgriBench-VL-4K, achieving higher pairwise win rates in the LLM-as-a-judge evaluation. Meanwhile, it remains competitive on the text-only AgriBench-13K with no noticeable degradation of language ability. Ablation studies further confirm consistent gains from our alignment and GRPO refinement stages. We will open source all of the resources to support reproducible research and deployment in low-resource agricultural settings.
Problem

Research questions and friction points this paper is trying to address.

Addresses scarcity of agricultural vision-language models and datasets
Develops specialized multimodal framework for agricultural applications
Creates rigorous evaluation suite for agricultural vision-language understanding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Largest agricultural vision-language corpus via multi-agent generation
Specialized model trained with progressive curriculum learning
Compact evaluation suite with multi-metric LLM-as-a-judge framework
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
B
Bo Yang
College of Computer Science and Technology, Zhejiang University
Y
Yunkui Chen
College of Software Technology, Zhejiang University
L
Lanfei Feng
College of Software Technology, Zhejiang University
Y
Yu Zhang
College of Computer Science and Technology, Zhejiang University
X
Xiao Xu
College of Computer Science and Technology, Zhejiang University
J
Jianyu Zhang
College of Computer Science and Technology, Zhejiang University
N
Nueraili Aierken
College of Computer Science and Technology, Zhejiang University
Runhe Huang
Runhe Huang
Professor, Faculty of Computer and Information Science, Hosei University
AIbrain modelingmachine intelligencecognitive computing
H
Hongjian Lin
College of Biosystems Engineering and Food Science, Zhejiang University
Y
Yibin Ying
College of Biosystems Engineering and Food Science, Zhejiang University
Shijian Li
Shijian Li
zhejiang university
pervasive computinghuman computer interactionartificial intelligence