🤖 AI Summary
This study addresses the silent failures in single-image nutrition estimation caused by undetected food items and regions that inadequately support portion size inference. To overcome these limitations, this work proposes a fine-tuning-free verification and recovery framework that leverages multimodal large language models combined with retrieval-augmented generation. The approach effectively decouples item identification, region validity verification, and gap repair, employing directed recovery to fill omissions and consolidating them into a complete food set for accurate estimation. Notably, the proposed method operates entirely without ground-truth annotations during inference. Experimental results demonstrate significant improvements in item-level localization accuracy and recall, alongside substantial enhancements in both mass and energy estimation precision.
📝 Abstract
Single-image nutrition estimation can fail silently when visible foods are missed. Even when a food is correctly identified, its proposed region may not support portion estimation. We propose a framework that uses multimodal large language models (MLLMs) to inventory visible foods and separately verify food identity and whether each proposed 2D region supports portion estimation. One whole-image review uses these verification results to identify unresolved gaps and omitted foods, triggering at most one targeted recovery pass. Recovered regions are re-verified without access to the recovery prompt, then reconciled into a final item set for nutrition estimation. The framework requires no task-specific fine-tuning. Matched evaluation on common valid-output samples shows that item-level grounding improves mass accuracy across all tested settings and energy accuracy relative to an adapted retrieval baseline, with item-identity precision and recall also improving, while post-recovery visual coverage is assessed separately at inference time without ground-truth annotations.