SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing

📅 2026-07-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of language-guided segmentation in drone videos, where targets are often small, visually ambiguous, and captured under dynamic viewpoints. To tackle these issues, the authors propose SkyVLaM, a novel model that introduces a temporal basis-aware perceptron to construct sparse visual tokens. By regularizing sparse bases to fuse complementary temporal cues, the method adaptively selects temporally consistent dense segments for high-resolution analysis, enabling query-guided segmentation in conjunction with a large language model. The core innovation lies in the first adaptive sparse-to-dense visual token mechanism, which optimizes token budget allocation. Additionally, the authors introduce SkyVid, the first multimodal understanding and pixel-level annotated dataset specifically designed for drone video. Experiments demonstrate that the proposed approach significantly advances language-guided video segmentation performance, validating the efficacy of its efficient visual token allocation strategy.
📝 Abstract
Recent advances in Multimodal Large Language Models (MLLMs) have significantly improved remote sensing (RS) multimodal understanding. Language-conditioned segmentation is crucial for fine-grained target understanding in Unmanned Aerial Vehicle (UAV) videos. However, this task remains challenging due to the prevalence of small, visually ambiguous targets and dynamic aerial perspectives. In this paper, we propose SkyVLaM, a multimodal large language model for UAV video understanding. SkyVLaM constructs sparse tokens directly from patch-level video representations through a temporal basis perceiver, regularizes the sparse basis to encourage complementary temporal cues, and adaptively selects a temporally coherent dense segment for high-resolution inspection. The resulting sparse and dense tokens are jointly processed by a large language model for query-conditioned segmentation. We further build SkyVid, consisting of SkyVid-VGCG and SkyVid-RVOS for video grounded conversation generation and referring video object segmentation, respectively. SkyVid contains 101 videos, 33.6K frames, and 1.53M pixel-level object instances. Experiments show that SkyVLaM provides a more effective allocation of the visual token budget and improves language-conditioned video segmentation in UAV scenarios.
Problem

Research questions and friction points this paper is trying to address.

UAV video understanding
language-conditioned segmentation
small and ambiguous targets
dynamic aerial perspectives
remote sensing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Large Language Model
UAV Video Understanding
Language-conditioned Segmentation
Temporal Basis Perceiver
Sparse-Dense Token Fusion
🔎 Similar Papers
2024-06-09Annual Meeting of the Association for Computational LinguisticsCitations: 13