Fine-Grained Food Image Understanding via Target-Aware Data Alignment

๐Ÿ“… 2026-07-28
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses domain shift and misalignment between visual and textual modalities in web-crawled food images by proposing a target-aware data alignment framework. The approach integrates vision-language model (VLM)-driven refinement of imageโ€“caption pairs, a CLIP-style multi-expert retrieval model, and a hierarchical decision fusion strategy to construct high-quality training sets while significantly enhancing fine-grained food recognition and description performance. Caption refinement alone yields approximately a 19% improvement in retrieval accuracy, and the full method more than doubles the retrieval score compared to a pure VLM baseline while achieving lower inference overhead.
๐Ÿ“ Abstract
Fine-grained food visual--semantic understanding requires models to capture subtle distinctions across ingredients, cooking methods, doneness, color, texture, and plate composition. Although CLIP-style vision-language models provide a natural framework for this task, their effectiveness is limited when training relies on heterogeneous web-collected image--text pairs. Such data often exhibit a web-to-target domain gap and cross-modal misalignment, where images differ from the target distribution and captions are noisy, multilingual, or weakly grounded in visual content. We propose a data-centric multimodal alignment method for fine-grained food description and recognition. Our method first performs target-aware data selection to identify visually relevant training subsets, then applies VLM-based caption refinement to generate visually grounded, target-style descriptions. Using these curated image--caption pairs, we train complementary CLIP-style retrieval experts and further combine their decisions through a hierarchical VLM-assisted multi-expert decision-level fusion strategy that invokes the VLM only when experts disagree. Experiments show that our data refinement strategy significantly improves retrieval performance over naive web supervision, with VLM-based caption refinement alone yielding an average performance gain of approximately 19%. Our full method also achieves more than twice the retrieval score of pure VLM-based retrieval while remaining substantially more efficient.
Problem

Research questions and friction points this paper is trying to address.

fine-grained food understanding
domain gap
cross-modal misalignment
web-collected data
visual-semantic alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

target-aware data selection
VLM-based caption refinement
multimodal alignment
multi-expert fusion
fine-grained food understanding
๐Ÿ”Ž Similar Papers
No similar papers found.
J
Jui-Feng Chi
Elmore Family School of Electrical and Computer Engineering, Purdue University, West Lafayette, IN, USA
W
Wei-Lun Chu
Elmore Family School of Electrical and Computer Engineering, Purdue University, West Lafayette, IN, USA
B
Bruce Coburn
Elmore Family School of Electrical and Computer Engineering, Purdue University, West Lafayette, IN, USA
Jinge Ma
Jinge Ma
Nanjing Institute of Geography and Limnology, Chinese Academy of Sciences
algal bloomremote sensinglake ecosystem
F
Fengqing Zhu
Elmore Family School of Electrical and Computer Engineering, Purdue University, West Lafayette, IN, USA