Hallucination Behavior in Multimodal LLMs Across Agricultural Image Interpretation and Generation Tasks

๐Ÿ“… 2026-05-26
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the prevalence of hallucinations in multimodal large language models (MLLMs) when applied to agricultural image understanding and generationโ€”errors that often violate biological or agronomic principles and risk misleading decision-making. The work presents the first systematic evaluation of such hallucinatory behavior across both image-to-text interpretation and text-to-image generation tasks, introducing a fine-grained assessment framework tailored to agriculture that encompasses biological consistency, contextual accuracy, and agronomic plausibility. Leveraging mainstream models including Gemini, GPT-5, and LLaVA, the authors conduct multimodal evaluations under zero-shot and few-shot settings, complemented by human validation. Results reveal that zero-shot image interpretation achieves only 63%โ€“75% accuracy, improving to 86.8% with few-shot learning, while 91% of generated images exhibit biological inconsistencies, underscoring significant reliability limitations and high hallucination risks in current MLLMs for agricultural applications.
๐Ÿ“ Abstract
Large Language Models (LLMs) are being rapidly adopted in agricultural imaging applications, ranging from crop interpretation to synthetic field image generation. However, these models frequently exhibit hallucinations outputs that appear confident yet deviate from biological or environmental reality potentially leading to misinformed agronomic insights. This study investigates such hallucinations in two complementary directions: image-to-text, where LLMs interpret crop or field imagery to describe conditions such as biotic and abiotic stresses, and text-to-image, where models generate synthetic agricultural scenes based on descriptive prompts. We examine errors involving biological inconsistency, contextual inaccuracy, and agronomic implausibility, evaluating the outputs under domain-informed criteria across multiple imaging modalities. Our analysis identifies recurring hallucination patterns within both interpretive and generative tasks. In image interpretation, LLMs (e.g., Gemma, LLAVA, Qwen, and MiniCPM) achieved modest zero-shot accuracy (63 to 75 percent), whereas few-shot prompting improved performance up to 86.8 percent, exhibiting false detections and missed infections, indicating residual hallucination effects. In text-to-image tasks, advanced models such as GPT-5 and Gemini 2.5 Flash generate up to 91 percent biologically inconsistent scenes under relaxed prompt constraints, revealing fundamental weaknesses in current LLMs. This systematic assessment of visual reasoning and generation offers critical insights toward enhancing the reliability and trustworthiness of LLM-based agricultural imaging platforms.
Problem

Research questions and friction points this paper is trying to address.

hallucination
multimodal LLMs
agricultural imaging
image interpretation
image generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

multimodal hallucination
agricultural imaging
biological consistency
few-shot prompting
synthetic image generation
๐Ÿ”Ž Similar Papers
๐Ÿ’ผ Related Jobs
No related jobs found.
Partho Ghose
Partho Ghose
Graduate Research Assistant at Texas A&M University
Artificial IntelligenceComputer VisionPublic HealthAgriculture Engineering
A
Al Bashir
Department of Biological & Agricultural Engineering, Texas A&M AgriLife Research, Texas A&M University System, Dallas, TX 75252, USA.
P
Prem Raj
Department of Biological & Agricultural Engineering, Texas A&M AgriLife Research, Texas A&M University System, Dallas, TX 75252, USA.
A
Azlan Zahid
Department of Biological & Agricultural Engineering, Texas A&M AgriLife Research, Texas A&M University System, Dallas, TX 75252, USA.