Improving Image-Based Nutrition Estimation Through Multimodal Food-Item Verification and Recovery

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the silent failures in single-image nutrition estimation caused by undetected food items and regions that inadequately support portion size inference. To overcome these limitations, this work proposes a fine-tuning-free verification and recovery framework that leverages multimodal large language models combined with retrieval-augmented generation. The approach effectively decouples item identification, region validity verification, and gap repair, employing directed recovery to fill omissions and consolidating them into a complete food set for accurate estimation. Notably, the proposed method operates entirely without ground-truth annotations during inference. Experimental results demonstrate significant improvements in item-level localization accuracy and recall, alongside substantial enhancements in both mass and energy estimation precision.
📝 Abstract
Single-image nutrition estimation can fail silently when visible foods are missed. Even when a food is correctly identified, its proposed region may not support portion estimation. We propose a framework that uses multimodal large language models (MLLMs) to inventory visible foods and separately verify food identity and whether each proposed 2D region supports portion estimation. One whole-image review uses these verification results to identify unresolved gaps and omitted foods, triggering at most one targeted recovery pass. Recovered regions are re-verified without access to the recovery prompt, then reconciled into a final item set for nutrition estimation. The framework requires no task-specific fine-tuning. Matched evaluation on common valid-output samples shows that item-level grounding improves mass accuracy across all tested settings and energy accuracy relative to an adapted retrieval baseline, with item-identity precision and recall also improving, while post-recovery visual coverage is assessed separately at inference time without ground-truth annotations.
Problem

Research questions and friction points this paper is trying to address.

nutrition estimation
food detection
portion estimation
multimodal verification
silent failure
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Large Language Models
Nutrition Estimation
Food-Item Verification
Targeted Recovery
Zero Fine-tuning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jingbo Yue
Elmore Family School of Electrical and Computer Engineering, Purdue University
B
Bruce Coburn
Elmore Family School of Electrical and Computer Engineering, Purdue University
Jinge Ma
Jinge Ma
Nanjing Institute of Geography and Limnology, Chinese Academy of Sciences
algal bloomremote sensinglake ecosystem
J
Jui-Feng Chi
Elmore Family School of Electrical and Computer Engineering, Purdue University
F
Fengqing Zhu
Elmore Family School of Electrical and Computer Engineering, Purdue University