The Role of Natural Language Understanding in Multimodal Video-Based Dengue Diagnosis

📅 2026-08-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant challenges in identifying dengue virus–infected mosquito behaviors, which arise from their small size, rapid movement, and environmental interference. The work proposes a novel multimodal framework that integrates YOLO-based object detection with CLIP contrastive learning, leveraging biological text prompts to align visual and semantic information within a shared embedding space. Supervised bidirectional contrastive fine-tuning and temporal aggregation strategies are employed to classify infection status effectively. The method achieves 98.54% frame-level accuracy and 99.91% sensitivity, with perfect video-level prediction performance. Ablation studies confirm the critical contributions of CLIP’s semantic alignment and the fine-tuning mechanism, underscoring the necessity and efficacy of multimodal representation learning in mosquito-borne virus behavioral analysis.
📝 Abstract
Detecting infection-related behavioral changes in mosquitoes from video data is challenging because mosquitoes are small, move rapidly and irregularly, and are affected by environmental factors such as background, lighting, and shadows, which can make reliable feature extraction difficult. In this study, a YOLO- and Contrastive Language-Image Pre-training (CLIP)-based vision-language framework is proposed to classify mosquito flight frames of uninfected and Dengue virus serotype 2 (DENV2)-infected mosquitoes. First, YOLO is used to isolate mosquito regions from the background. Then, visual features extracted from video frames are aligned with biologically meaningful textual prompts in a shared embedding space. The multimodal model was fine-tuned using supervised bidirectional contrastive learning and evaluated through frame-level image-text similarity-based classification. The results show that the proposed method achieved 98.54% accuracy and 99.91% sensitivity at the frame level. After temporal aggregation of frame-level information, the model achieved complete video-level performance. The ablation results showed that fine-tuning and CLIP-based representations were essential for this domain, while the textual branch provided semantic image-text alignment rather than an accuracy advantage over the vision-only model. These findings suggest that vision-language models can provide a useful framework for analyzing infection-related biological behaviors from video data.
Problem

Research questions and friction points this paper is trying to address.

multimodal video analysis
dengue diagnosis
mosquito behavior
natural language understanding
feature extraction
Innovation

Methods, ideas, or system contributions that make the work stand out.

vision-language model
CLIP
YOLO
contrastive learning
multimodal video analysis
🔎 Similar Papers
2024-06-09Annual Meeting of the Association for Computational LinguisticsCitations: 13
💼 Related Jobs
No related jobs found.
D
Danial Sharifrazi
Institute for Intelligent Systems Research and Innovations (IISRI), Deakin University, Geelong, Australia
S
Saadat Behzadi
Department of Electronic Engineering, University of Bologna, Bologna, Italy
J
Julakha Jahan Jui
Institute for Intelligent Systems Research and Innovations (IISRI), Deakin University, Geelong, Australia
M
Mojtaba Mohammadi
Institute for Intelligent Systems Research and Innovations (IISRI), Deakin University, Geelong, Australia
N
Nouman Javed
Institute for Intelligent Systems Research and Innovations (IISRI), Deakin University, Geelong, Australia
R
Roohallah Alizadehsani
Institute for Intelligent Systems Research and Innovations (IISRI), Deakin University, Geelong, Australia
P
Prasad N. Paradkar
CSIRO Health and Biosecurity, Australian Animal Health Laboratory, Geelong, Australia
Asim Bhatti
Asim Bhatti
Professor; Deakin University
Neural and Cognitive SystemsBrain-on-a-ChipNeuroengineering