Learning to Understand Body Language from Flight through Robust 3D Avatar Placing

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of long-range human action and intention recognition, which is hindered by a scarcity of real-world data and consequently limits the development of socially intelligent drones. To overcome this, the authors introduce the Drones2BodyLanguage dataset, embedding 3D virtual avatars expressing ten communicative intentions into real 4K aerial videos with precise position, scale, and orientation. They propose a lightweight geometric world model that combines semantic anchors with monocular depth flow to predict placement via an affine combination using rigid-body-invariant weights, and employ SVD-based ground-plane rotation fitting to achieve temporally stable and photorealistic re-rendering across frames. Experiments demonstrate that models trained solely on this synthetic data significantly improve intention recognition accuracy on real, retargeted, and generated actions under scene–action disjoint test splits, with strong generalization validated in two in-the-wild scenarios.
📝 Abstract
Perceiving human motion and intent at long range is a prerequisite for socially intelligent aerial robots, yet the data to learn it barely exists. We introduce Drones2BodyLanguage, a dataset grounding human motion in real UAV footage: avatars manifesting ten communicative intents are placed into unmodified 4K drone scenes with metrically correct position, scale and orientation, maintained over hundreds of frames of camera motion. Enabling it is a lightweight geometric world model of the local scene - semantically selected anchors lifted to 3D through streaming monocular depth - in which a placement point is predicted as an affine anchor combination with provably rigid-invariant weights, and re-rendered under an SVD-fitted ground rotation. Across twelve architectures on scene- and motion-disjoint splits, training on placed data lifts mean intent accuracy by a wide margin for real, retargeted and generated motion alike, with gains confirmed on two in-the-wild scenes.
Problem

Research questions and friction points this paper is trying to address.

human motion understanding
aerial robots
body language
long-range perception
communicative intent
Innovation

Methods, ideas, or system contributions that make the work stand out.

3D avatar placement
monocular depth estimation
rigid-invariant affine anchoring
UAV-based human intent perception
geometric world modeling
🔎 Similar Papers
No similar papers found.