NutriVision: Ingredient-Conditioned Fusion and Prediction for Single-Image Food Nutrition Estimation

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of missing depth information and insufficient multimodal fusion in nutrition estimation from a single RGB image by proposing a framework that integrates visual geometry with ingredient semantics. Methodologically, DepthAnything-V3 is employed to extract geometric cues, while CLIP encodes semantic features. An Ingredient-Conditioned Frequency Alignment Fusion Module (IC-FAFM) is designed to facilitate efficient cross-modal integration, complemented by an Ingredient-Aware Mask Prediction Head (IA-MPH) for internal semantic modeling and conditioned nutrient prediction. This approach enables accurate estimation of calories and other nutrients without requiring specialized hardware. Evaluated on the Nutrition5k dataset, the proposed method achieves a Percentage Mean Absolute Error (PMAE) of 13.60%, significantly outperforming existing baselines and demonstrating its effectiveness.
📝 Abstract
Nutrition estimation is a fundamental task in consumer diet tracking, clinical dietetics, chronic disease management, sports and hospital nutrition, and broader food computing systems. The existing approaches have progressed along two largely separate axes, vision models that rely on calibrated RGB-depth captures and ingredient-aware methods that use textual cues but use limited multimodal fusion. We introduce NutriVision, an end-to-end framework that leverages visual geometry and ingredient semantics to estimate calories, mass, fat content, carbohydrates, and protein from a single RGB image and an optional ingredient list. It obtains the unavailable depth modality using DepthAnything-V3 and encodes ingredient descriptions using CLIP. It integrates three complementary mechanisms: (1) an \emph{Ingredient-Conditioned Frequency-Aligned Fusion Module (IC-FAFM)}, which uses textual guidance to reweight and align RGB-depth frequency components; (2) an \emph{Ingredient-Aware Mask-based Prediction Head (IA-MPH)}, whose gating and channel masks are conditioned on food identity; and (3) modality-specific \emph{Internal Semantic Modeling (ISM)} blocks. On the Nutrition5k dataset, NutriVision achieves a mean PMAE of $\mathbf{13.60\pm0.10\%}$, outperforming our IGSMNet implementation by $0.89$ percentage points and OmniFood8k by $2.90$ percentage points (both $p<0.001$). The module-level ablations identify the ingredient-aware prediction head as the primary architectural contributor, improving mean PMAE by $1.50\pm0.17$ percentage points ($p<0.001$). These results demonstrate that ingredient-conditioned prediction and frequency-aware RGB-depth fusion provide measurable gains for single-image nutrient estimation. More broadly, NutriVision offers a practical route toward nutrition-assessment systems that exploit geometric and semantic cues without requiring specialized depth-sensing hardware
Problem

Research questions and friction points this paper is trying to address.

food nutrition estimation
single RGB image
multimodal fusion
ingredient semantics
depth modality
Innovation

Methods, ideas, or system contributions that make the work stand out.

Ingredient-Conditioned Fusion
Single-Image Nutrition Estimation
Frequency-Aligned Fusion
Monocular Depth Estimation
Multimodal Learning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Aman Kumar
Indraprastha Institute of Information Technology Delhi (IIIT-Delhi), New Delhi, India
A
Avinash Anand
Singapore Institute of Technology, Singapore
C
Chaitanya Lakhchaura
Indraprastha Institute of Information Technology Delhi (IIIT-Delhi), New Delhi, India
A
Ashutosh Kumar
Rochester Institute of Technology, Rochester, NY, USA
Akshita Abrol
Akshita Abrol
Singapore Institute of Technology
T
Timothy Liu
NVIDIA AI Technology Centre, Singapore
Z
Zhengkui Wang
Singapore Institute of Technology, Singapore
R
Rajiv Ratn Shah
Indian Institute of Technology Kanpur, Kanpur, India