Prompt-to-Gesture: Measuring the Capabilities of Image-to-Video Deictic Gesture Generation

๐Ÿ“… 2026-04-16
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the scarcity and high acquisition cost of real-world hand gesture data, which has long hindered progress in gesture recognition research. It introduces, for the first time, an image-to-video foundation model for zero-shot deictic gesture synthesis, capable of generating high-fidelity, diverse, and realistic gesture videos from only a few reference images of a person and natural language prompts. The synthesized dataset effectively compensates for the lack of real data, and when used in mixed training with limited real samples, it substantially enhances the performance of downstream deep learning models. This demonstrates the methodโ€™s superior data efficiency and generation quality, offering a promising avenue for scalable and cost-effective gesture data creation.

Technology Category

Computer Vision: Biometrics, Face, Gesture & PoseNatural Language Processing: Code Generation / Program Synthesis from Natural LanguageMachine Learning: Deep Generative Models & Autoencoders

Application Category

Economics, Online Markets and Human Computation: Data quality aspects of human-annotated datasetsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsWeb Mining and Content Analysis: Large pretrained models with web data
๐Ÿ“ Abstract
Gesture recognition research, unlike NLP, continues to face acute data scarcity, with progress constrained by the need for costly human recordings or image processing approaches that cannot generate authentic variability in the gestures themselves. Recent advancements in image-to-video foundation models have enabled the generation of photorealistic, semantically rich videos guided by natural language. These capabilities open up new possibilities for creating effort-free synthetic data, raising the critical question of whether video Generative AI models can augment and complement traditional human-generated gesture data. In this paper, we introduce and analyze prompt-based video generation to construct a realistic deictic gestures dataset and rigorously evaluate its effectiveness for downstream tasks. We propose a data generation pipeline that produces deictic gestures from a small number of reference samples collected from human participants, providing an accessible approach that can be leveraged both within and beyond the machine learning community. Our results demonstrate that the synthetic gestures not only align closely with real ones in terms of visual fidelity but also introduce meaningful variability and novelty that enrich the original data, further supported by superior performance of various deep models using a mixed dataset. These findings highlight that image-to-video techniques, even in their early stages, offer a powerful zero-shot approach to gesture synthesis with clear benefits for downstream tasks.
Problem

Research questions and friction points this paper is trying to address.

gesture recognition
data scarcity
deictic gestures
synthetic data
image-to-video generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

image-to-video generation
deictic gesture synthesis
prompt-based generation
synthetic gesture dataset
zero-shot learning
๐Ÿ”Ž Similar Papers
No similar papers found.
Hassan Ali
Hassan Ali
University of Hamburg
D
Doreen Jirak
Behavioral Lab, Department of Product Development, University of Antwerp, Belgium
L
Luca Mรผller
Knowledge Technology Group, Department of Informatics, University of Hamburg, Germany
S
Stefan Wermter
Knowledge Technology Group, Department of Informatics, University of Hamburg, Germany