🤖 AI Summary
This study addresses the challenges of missing depth information and insufficient multimodal fusion in nutrition estimation from a single RGB image by proposing a framework that integrates visual geometry with ingredient semantics. Methodologically, DepthAnything-V3 is employed to extract geometric cues, while CLIP encodes semantic features. An Ingredient-Conditioned Frequency Alignment Fusion Module (IC-FAFM) is designed to facilitate efficient cross-modal integration, complemented by an Ingredient-Aware Mask Prediction Head (IA-MPH) for internal semantic modeling and conditioned nutrient prediction. This approach enables accurate estimation of calories and other nutrients without requiring specialized hardware. Evaluated on the Nutrition5k dataset, the proposed method achieves a Percentage Mean Absolute Error (PMAE) of 13.60%, significantly outperforming existing baselines and demonstrating its effectiveness.
📝 Abstract
Nutrition estimation is a fundamental task in consumer diet tracking, clinical dietetics, chronic disease management, sports and hospital nutrition, and broader food computing systems. The existing approaches have progressed along two largely separate axes, vision models that rely on calibrated RGB-depth captures and ingredient-aware methods that use textual cues but use limited multimodal fusion. We introduce NutriVision, an end-to-end framework that leverages visual geometry and ingredient semantics to estimate calories, mass, fat content, carbohydrates, and protein from a single RGB image and an optional ingredient list. It obtains the unavailable depth modality using DepthAnything-V3 and encodes ingredient descriptions using CLIP. It integrates three complementary mechanisms: (1) an \emph{Ingredient-Conditioned Frequency-Aligned Fusion Module (IC-FAFM)}, which uses textual guidance to reweight and align RGB-depth frequency components; (2) an \emph{Ingredient-Aware Mask-based Prediction Head (IA-MPH)}, whose gating and channel masks are conditioned on food identity; and (3) modality-specific \emph{Internal Semantic Modeling (ISM)} blocks. On the Nutrition5k dataset, NutriVision achieves a mean PMAE of $\mathbf{13.60\pm0.10\%}$, outperforming our IGSMNet implementation by $0.89$ percentage points and OmniFood8k by $2.90$ percentage points (both $p<0.001$). The module-level ablations identify the ingredient-aware prediction head as the primary architectural contributor, improving mean PMAE by $1.50\pm0.17$ percentage points ($p<0.001$). These results demonstrate that ingredient-conditioned prediction and frequency-aware RGB-depth fusion provide measurable gains for single-image nutrient estimation. More broadly, NutriVision offers a practical route toward nutrition-assessment systems that exploit geometric and semantic cues without requiring specialized depth-sensing hardware