MAVR-Net: Robust Multi-View Learning for MAV Action Recognition with Cross-View Attention

📅 2025-10-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the insufficient robustness and accuracy of micro-air-vehicle (MAV) action recognition in complex aerial scenes, this paper proposes MAVR-Net, a multi-view fusion network. MAVR-Net jointly models three complementary modalities—RGB frames, optical flow fields, and semantic segmentation masks—leveraging a ResNet backbone, a multi-scale feature pyramid, and a cross-view attention mechanism to enable deep spatiotemporal feature interaction. A novel multi-view alignment loss is introduced to explicitly enforce semantic consistency across modalities. End-to-end trained, MAVR-Net achieves state-of-the-art performance on multiple MAV action benchmarks: 97.8%, 96.5%, and 92.8% accuracy for short-, medium-, and long-range action recognition, respectively. The framework provides a scalable, highly robust multimodal learning architecture for vision-based MAV behavioral understanding.

Technology Category

Computer Vision: Multi-modal VisionIntelligent Robots: Multimodal Perception & Sensor FusionMachine Learning: Multi-instance/Multi-view Learning

Application Category

Responsible Web: Machine-in-the-loop, human agency and autonomySearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Recognizing the motion of Micro Aerial Vehicles (MAVs) is crucial for enabling cooperative perception and control in autonomous aerial swarms. Yet, vision-based recognition models relying only on RGB data often fail to capture the complex spatial temporal characteristics of MAV motion, which limits their ability to distinguish different actions. To overcome this problem, this paper presents MAVR-Net, a multi-view learning-based MAV action recognition framework. Unlike traditional single-view methods, the proposed approach combines three complementary types of data, including raw RGB frames, optical flow, and segmentation masks, to improve the robustness and accuracy of MAV motion recognition. Specifically, ResNet-based encoders are used to extract discriminative features from each view, and a multi-scale feature pyramid is adopted to preserve the spatiotemporal details of MAV motion patterns. To enhance the interaction between different views, a cross-view attention module is introduced to model the dependencies among various modalities and feature scales. In addition, a multi-view alignment loss is designed to ensure semantic consistency and strengthen cross-view feature representations. Experimental results on benchmark MAV action datasets show that our method clearly outperforms existing approaches, achieving 97.8%, 96.5%, and 92.8% accuracy on the Short MAV, Medium MAV, and Long MAV datasets, respectively.
Problem

Research questions and friction points this paper is trying to address.

Recognizing MAV motion for cooperative perception in autonomous aerial swarms.
Overcoming RGB-only limitations in capturing complex spatiotemporal MAV characteristics.
Enhancing action recognition robustness through multi-view data fusion.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-view learning combines RGB, optical flow, segmentation masks
Cross-view attention module models dependencies between modalities
Multi-scale feature pyramid preserves spatiotemporal motion details
🔎 Similar Papers
2024-07-16arXiv.orgCitations: 3
💼 Related Jobs
No related jobs found.
Universiti Sains Malaysia
N
Nengbo Zhang
School of Aerospace Engineering, Universiti Sains Malaysia, 14300 Nibong Tebal, Pulau Pinang, Malaysia.
Hann Woei Ho
Hann Woei Ho
Senior Lecturer/ Assistant Professor, School of Aerospace Engineering, Universiti Sains Malaysia
Unmanned Aerial VehiclesControl TheoryComputer VisionMachine Learning