🤖 AI Summary
This study addresses the challenge of accurately distinguishing dengue virus–infected mosquitoes from control specimens in complex video backgrounds, where conventional methods struggle to extract discriminative features. To overcome this, the authors propose a three-stage framework: first, YOLOv11M detects and crops mosquito regions to eliminate background interference; second, a Vision Transformer (ViT) extracts spatial features with strong representational capacity; and third, a ConvGRU models long-term temporal dynamics, enabling joint optimization of spatial and temporal information. The proposed integration of ViT and ConvGRU—selected after comparative evaluation against various RNNs and their convolutional variants—achieves state-of-the-art performance, yielding 88.88% accuracy, 84.45% precision, 82.82% recall, and 82.81% F1 score, thereby significantly enhancing the reliability of mosquito-borne disease detection in challenging visual environments.
📝 Abstract
Identifying dengue virus-infected mosquitoes from control mosquitoes is a major challenge in analyzing mosquito locomotion behavior due to the small size and complexity of the video background. Conventional AI methods are often unable to extract accurate features from video frames and produce erroneous features. In this study, a three-step framework is introduced: first, mosquitoes are identified and the background is removed using the YOLO 11M model, then visual features are extracted using the Vision Transformer (ViT), and finally the videos are classified with a convolutional GRU (ConvGRU) classifier. A comparative analysis of different models, including Recurrent Neural Network (RNN), Long Short-Term Memory (LSTM), Gated Recurrent Unit (GRU), and their convolutional versions showed that the ConvGRU model achieved the best performance; it achieved 88.88% accuracy, 84.45% precision, 82.82% recall, and 82.81% F1 score. These results demonstrate that combining convolutional models with sequence-based networks, especially in the ConvGRU model, allows the simultaneous extraction of precise spatial features and long-term temporal dependencies from mosquito movements. Finally, the proposed framework provides a reliable solution for analyzing mosquito behavior in complex environments.