🤖 AI Summary
To address the insufficient robustness and accuracy of micro-air-vehicle (MAV) action recognition in complex aerial scenes, this paper proposes MAVR-Net, a multi-view fusion network. MAVR-Net jointly models three complementary modalities—RGB frames, optical flow fields, and semantic segmentation masks—leveraging a ResNet backbone, a multi-scale feature pyramid, and a cross-view attention mechanism to enable deep spatiotemporal feature interaction. A novel multi-view alignment loss is introduced to explicitly enforce semantic consistency across modalities. End-to-end trained, MAVR-Net achieves state-of-the-art performance on multiple MAV action benchmarks: 97.8%, 96.5%, and 92.8% accuracy for short-, medium-, and long-range action recognition, respectively. The framework provides a scalable, highly robust multimodal learning architecture for vision-based MAV behavioral understanding.
📝 Abstract
Recognizing the motion of Micro Aerial Vehicles (MAVs) is crucial for enabling cooperative perception and control in autonomous aerial swarms. Yet, vision-based recognition models relying only on RGB data often fail to capture the complex spatial temporal characteristics of MAV motion, which limits their ability to distinguish different actions. To overcome this problem, this paper presents MAVR-Net, a multi-view learning-based MAV action recognition framework. Unlike traditional single-view methods, the proposed approach combines three complementary types of data, including raw RGB frames, optical flow, and segmentation masks, to improve the robustness and accuracy of MAV motion recognition. Specifically, ResNet-based encoders are used to extract discriminative features from each view, and a multi-scale feature pyramid is adopted to preserve the spatiotemporal details of MAV motion patterns. To enhance the interaction between different views, a cross-view attention module is introduced to model the dependencies among various modalities and feature scales. In addition, a multi-view alignment loss is designed to ensure semantic consistency and strengthen cross-view feature representations. Experimental results on benchmark MAV action datasets show that our method clearly outperforms existing approaches, achieving 97.8%, 96.5%, and 92.8% accuracy on the Short MAV, Medium MAV, and Long MAV datasets, respectively.