🤖 AI Summary
This work addresses the challenges of multimodal fusion and efficient retrieval in vision-language fact-checking by proposing a lightweight, modularly decoupled dual-retriever RAG architecture. The framework integrates similarity-based textual retrieval with API-driven reverse image search (RIS), enabling claim verification through a single invocation of a multimodal large language model (GPT-5.1). This approach substantially reduces inference costs—averaging just \$0.013 per query—while achieving third place in the AVerImaTeC shared task. The method demonstrates high performance, strong reproducibility, and considerable optimization potential, establishing a cost-effective and scalable baseline for multimodal fact-checking.
📝 Abstract
In this paper, we present our 3rd place system in the AVerImaTeC shared task, which combines our last year's retrieval-augmented generation (RAG) pipeline with a reverse image search (RIS) module. Despite its simplicity, our system delivers competitive performance with a single multimodal LLM call per fact-check at just $0.013 on average using GPT5.1 via OpenAI Batch API. Our system is also easy to reproduce and tweak, consisting of only three decoupled modules - a textual retrieval module based on similarity search, an image retrieval module based on API-accessed RIS, and a generation module using GPT5.1 - which is why we suggest it as an accesible starting point for further experimentation. We publish its code and prompts, as well as our vector stores and insights into the scheme's running costs and directions for further improvement.