MultiFly: A Real-World Multimodal Aerial Dataset with Annotation-Efficient Label Transfer and Cross-Modal Semantic Consistency

๐Ÿ“… 2026-10-07
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the high annotation costs and cross-modal semantic inconsistencies in multimodal perception for low-altitude UAVs by constructing the first synchronized public dataset integrating RGB, thermal infrared, LiDAR, and radar modalities. By leveraging shared geometric representations and frame-level consistent annotations, this work proposes a cross-modal automatic label propagation method guided by minimal human priors. Combined with multi-sensor fusion and semantic consistency verification techniques, the approach enables highly efficient data annotation. The project generates large-scale, high-quality labeled data, achieving an average label consistency of 89.93% and cross-modal semantic consistency exceeding 90%. Furthermore, it establishes a real-world four-modal semantic segmentation benchmark that significantly reduces annotation overhead while ensuring robust cross-modal alignment.
๐Ÿ“ Abstract
We introduce MultiFly, a real-world, low-altitude UAV dataset for semantic perception across RGB, thermal, LiDAR, and radar modalities. MultiFly provides 17,272 synchronized samples from four suburban scenes with frame-wise annotations for 15 semantic classes, together with calibration and GNSS-RTK/IMU measurements. To avoid costly and inconsistent modality-specific annotation, we propagate labels from only 115 manually annotated RGB images through shared geometric representations to all four modalities. This approach generates semantic labels for 17,157 additional RGB images, 17,272 thermal images, 840M LiDAR points, and 3.4M radar points. Transferred annotations achieve 89.93% average agreement with held-out manual annotations, and 90.94% average semantic consistency across all six modality pairs. We further establish semantic segmentation benchmarks for all four modalities, revealing distinct architectural behavior for dense LiDAR and sparse radar data. Taken together, MultiFly provides a scalable foundation for multimodal aerial perception and, to the best of our knowledge, the first public real-world low-altitude aerial benchmark that combines consistent frame-wise semantic annotations for RGB, thermal, LiDAR, and radar. Data at https://github.com/markus-42/multifly.
Problem

Research questions and friction points this paper is trying to address.

multimodal perception
UAV dataset
semantic segmentation
low-altitude aerial
annotation efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal aerial dataset
Annotation-efficient label transfer
Cross-modal semantic consistency
Semantic segmentation benchmark
Low-altitude UAV perception
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
M
Markus Gross
Autonomous Aerial Systems, Fraunhofer Institute IVI; Computer Vision Group, Technical University of Munich; Munich Center for Machine Learning (MCML); Computer Vision for Digital Twins, University of Cambridge
Andreas Greiner
Andreas Greiner
Lecturer, Faculty of Engineering, University of Freiburg
Simulation
T
Taehyoung Kim
Autonomous Aerial Systems, Fraunhofer Institute IVI
S
Sivasubiramaniam Subbiah
Autonomous Aerial Systems, Fraunhofer Institute IVI
T
Tomaลพ Cotiฤ
Autonomous Aerial Systems, Fraunhofer Institute IVI
S
Sai Bharadwaj Matha
Autonomous Aerial Systems, Fraunhofer Institute IVI; Institute of Innovative Mobility, Univ. of Applied Sciences Ingolstadt
C
Conrad Christoph
Autonomous Aerial Systems, Fraunhofer Institute IVI
Oussema Dhaouadi
Oussema Dhaouadi
PhD Student
Computer VisionDeep LearningRobotics
S
Simon Zieher
Autonomous Aerial Systems, Fraunhofer Institute IVI
S
Surya Vijaya Kumar
Autonomous Aerial Systems, Fraunhofer Institute IVI
G
Gordon Elger
Institute of Innovative Mobility, Univ. of Applied Sciences Ingolstadt
H
Henri MeeรŸ
Autonomous Aerial Systems, Fraunhofer Institute IVI
Olaf Wysocki
Olaf Wysocki
Assistant Research Professor, University of Cambridge
Computer VisionPhotogrammetryMachine Learning
P
Paul Spannaus
Autonomous Aerial Systems, Fraunhofer Institute IVI
Daniel Cremers
Daniel Cremers
Technical University of Munich
Computer VisionMachine LearningOptimizationRobotics