Bridging the Modality Gap in Forensic Image Retrieval

📅 2026-06-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of cross-modal semantic alignment in forensic image retrieval—spanning photographs, textual descriptions, hand-drawn sketches, and facial composites—by proposing the first unified retrieval framework based on multimodal large language models (MLLMs). The approach leverages MLLMs to generate structured textual descriptions, which are then combined with Sentence-BERT text embeddings and state-of-the-art visual features through a multimodal similarity fusion strategy to achieve robust cross-modal alignment. Experimental results demonstrate that the proposed paradigm significantly improves retrieval accuracy and robustness, particularly in scenarios with limited visual information or under noisy conditions. These findings underscore the framework’s practical utility and strong generalization capability in real-world forensic investigations.
📝 Abstract
Automated image retrieval plays an increasingly critical role in modern forensic analysis, supporting investigative workflows that rely on efficient comparison of visual evidence. While prior work has focused primarily on developing and optimizing multimodal retrieval systems, limited attention has been paid to evaluating the forensic applicability of these technologies across diverse real-world scenarios. In this study, we present a unified retrieval framework adapted to four key forensic tasks: (1) tattoo image retrieval given a tattoo query image; (2) tattoo retrieval guided by human-expert textual descriptions, modelling the common situation where a witness verbally describes a tattoo; (3) tattoo retrieval from hand-drawn sketches; and (4) face retrieval from forensic face sketches. Our system leverages a multimodal large language model (MLLM) to automatically generate structured textual descriptions for all queries and gallery images, followed by sentence-transformer embedding for text-based comparison. We evaluate retrieval using visual-only embeddings, text-only embeddings and a multimodal fusion strategy that combines text- and image-based similarity scores derived from state-of-the-art visual feature extractors relevant to each task. The fusion of modalities consistently improves retrieval precision and robustness, especially in scenarios where visual information is limited or noisy (e.g., sketches, partial tattoos, or fragmented witness statements). This work highlights the forensic value of a unified multimodal retrieval pipeline and demonstrates how modern MLLMs can operationalize challenging forensic tasks that traditionally rely on manual expert analysis. Our results position multimodal retrieval as a promising tool for supporting investigative workflows involving tattoos, facial composites, and witness descriptions.
Problem

Research questions and friction points this paper is trying to address.

forensic image retrieval
modality gap
multimodal retrieval
tattoo retrieval
face sketch
Innovation

Methods, ideas, or system contributions that make the work stand out.

multimodal retrieval
forensic image analysis
multimodal large language model
sketch-based retrieval
sentence-transformer embedding
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Ricardo González-Gazapo
Advanced Technologies Application Center (CENATAV), Havana, Cuba
A
Annette Morales-González
Advanced Technologies Application Center (CENATAV), Havana, Cuba
Yoanna Martínez-Díaz
Yoanna Martínez-Díaz
CENATAV
pattern recognitionimage processingface recognitioncomputer vision
Heydi Méndez-Vázquez
Heydi Méndez-Vázquez
CENATAV
Digital Image ProcessingBiometricsFace Recognition
M
Milton García-Borroto
Centro de Sistemas Complejos, Facultad de Física, Universidad de La Habana, Havana, Cuba