Hierarchical Multi-Modal Retrieval for Knowledge-Grounded News Image Captioning

📅 2026-06-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing image captioning methods struggle to generate context-rich descriptions that incorporate object attributes, event contexts, and deep semantic information. This work proposes a hierarchical multimodal article retrieval mechanism that leverages structure-aware text weighting and multidimensional similarity computation—encompassing content-visual, visual-visual, and discourse-position alignments—to accurately retrieve relevant articles from an external news knowledge base. By synergistically integrating a vision-language model (VLM) with a large language model (LLM), the framework generates news image captions enriched with contextual depth. The approach overcomes the limitations of conventional unimodal retrieval strategies, significantly enhancing both semantic richness and factual consistency. It achieved a competitive fifth place on the private test set of the ACM Multimedia EVENTA 2025 OpenEvent-V1 challenge, attaining a composite score of 0.2824.
📝 Abstract
Traditional image captioning methods often struggle to generate comprehensive, context-rich descriptions, especially for details not directly observable from visual cues. To overcome this, we propose a novel retrieval-augmented image captioning framework that generates captions with deeper insights, such as object attributes, event context, and underlying significance, by leveraging external knowledge. Our approach features a hierarchical multi-modal article retrieval mechanism that moves beyond monolithic text entities. This retrieval considers article structure-aware features, including weighted textual components (e.g., headlines, body sections) and visual placement patterns, alongside multi-faceted similarity computations (content--visual, visual--visual, and discourse positioning). A subsequent contextual relevance refinement stage further enhances the retrieved information. The retrieved articles then serve as the knowledge base for caption generation: first, a VLM generates a concise image description; second, we segment relevant information from the retrieved articles based on this description; and finally, an LLM utilizes both the description and extracted knowledge to generate a comprehensive, contextually detailed caption. We participated in the ACM Multimedia EVENTA 2025 Challenge and achieved 5th place with an overall score of 0.2824 on the private test set of the OpenEvent-V1 dataset. Source code is publicly released at https://github.com/mf0212/EVENTA-Challange.
Problem

Research questions and friction points this paper is trying to address.

image captioning
context-rich description
external knowledge
visual cues
knowledge grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

hierarchical multi-modal retrieval
retrieval-augmented captioning
structure-aware article features
knowledge-grounded image captioning
visual-language modeling
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Minh-Loi Nguyen
University of Science, VNU-HCM, Ho Chi Minh City, Vietnam; Vietnam National University, Ho Chi Minh City, Vietnam
X
Xuan-Vu Le
University of Science, VNU-HCM, Ho Chi Minh City, Vietnam; Vietnam National University, Ho Chi Minh City, Vietnam
L
Long-Bao Nguyen
University of Science, VNU-HCM, Ho Chi Minh City, Vietnam; Vietnam National University, Ho Chi Minh City, Vietnam
H
Hoang-Bach Ngo
University of Science, VNU-HCM, Ho Chi Minh City, Vietnam; Vietnam National University, Ho Chi Minh City, Vietnam
Trung-Nghia Le
Trung-Nghia Le
University of Science, VNU-HCM
Applied Deep LearningApplied Computer VisionMultimedia Security