OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diet

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of fine-grained food recognition, which is hindered by high intra-class variation and strong visual similarity among dishes, thereby limiting the accuracy of image-based dietary assessment. Building upon the PaliGemma-2-3B architecture, the authors employ parameter-efficient fine-tuning via LoRA and integrate three European dietary datasets into a unified dish vocabulary to train a vision-language model capable of multi-task instruction-based question answering. This approach achieves state-of-the-art performance in a domain-specific setting, surpassing large proprietary models such as Gemini, GPT, and Claude. In 3-fold cross-validation, the model attains 92.96% Top-1 dish recognition accuracy and 90.79% Exact-Set ingredient accuracy, significantly outperforming both CNN baselines and existing state-of-the-art methods.
📝 Abstract
Image based dietary assessment offers a scalable alternative to self reported food diaries, yet fine-grained food recognition remains challenging due to high intra-class variability and visually similar dishes. This study presents OliveGemma, a vision language model for recognising and reasoning about Mediterranean and European cuisine. Built on the open-weight PaliGemma-2-3B architecture, OliveGemma is fine-tuned with LoRA on a unified corpus of 17,340 images from three European research project datasets (MedGR, ODIN, and VIPPSTAR), reconciled into a vocabulary of 216 composed dish categories and paired with 102,642 instruction style question-answer items covering dish recognition, likely and visible ingredients, class boundary discrimination, visual evidence and overall visual food understanding. Under a 3-fold cross-validation scheme, OliveGemma achieves a top-1 accuracy of 92.96% +/- 0.91%, exceeding the strongest CNN baseline (DenseNet-121) by 7.31% and outperforming zero-shot frontier models with exact instructions and bounded classes including Gemini Flash 3 and 3.5, GPT-5.4 Mini, and Claude Haiku 4.6 by 8%, 46%, and 64% respectively. Furthermore, OliveGemma demonstrates competitive performance on Top-3 and Top-5 accuracy, being second best across CNNs and frontier models, surpassed only by DenseNet-121. In addition, OliveGemma achieves 90.79% +/- 1.3% Exact-Set on the likely ingredients of the food categories. These results demonstrate that PEFT adaptation of a small VLM can surpass substantially larger proprietary models on specialised food recognition. The model is publicly available at https://huggingface.co/JamesZar/OliveGemma-3B and the experiments and results can be found at https://github.com/tsiokris/OliveGemma.
Problem

Research questions and friction points this paper is trying to address.

food recognition
visual language model
Mediterranean diet
European cuisine
dietary assessment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Model
Parameter-Efficient Fine-Tuning
Dietary Assessment
Mediterranean Diet Recognition
LoRA
🔎 Similar Papers
No similar papers found.
D
Dimitrios I. Zaridis
Unit of Medical Technology & Intelligent Information Systems, University of Ioannina, Ioannina, Greece
T
Traianos Tsiokris
Unit of Medical Technology & Intelligent Information Systems, University of Ioannina, Ioannina, Greece
Vasileios C. Pezoulas
Vasileios C. Pezoulas
Unit of Medical Technology and Intelligent Information Systems, University of Ioannina & FORTH
D
Daphni Plati
Unit of Medical Technology & Intelligent Information Systems, University of Ioannina, Ioannina, Greece
E
Eugenia Mylona
Unit of Medical Technology & Intelligent Information Systems, University of Ioannina, Ioannina, Greece & Department of Medical Physics, School of Medicine, University of Patras, Patras, Greece
E
Eleni Georga
Unit of Medical Technology & Intelligent Information Systems, University of Ioannina, Ioannina, Greece
N
Nikos Tsiknakis
Computational BioMedicine Laboratory, Foundation for Research and Technology Hellas, Heraklion, Greece
Antonis Sakellarios
Antonis Sakellarios
Assistant professor of Biomedical Engineering, University of Patras
bioengineeringblood flow modelingatherosclerosisplaque growth
D
Dimitrios I. Fotiadis
Unit of Medical Technology & Intelligent Information Systems, University of Ioannina, Ioannina, Greece & Biomedical Research Institute, FORTH, Ioannina, Greece