UnrealPose: Leveraging Game Engine Kinematics for Large-Scale Synthetic Human Pose Data

📅 2026-01-02
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Acquiring large-scale, accurately annotated 3D human pose data in real-world scenarios is costly and scarce, while in-the-wild data often lacks ground-truth labels. To address this challenge, this work proposes UnrealPose-Gen, the first synthetic data generation framework built upon Unreal Engine 5’s offline rendering pipeline. The framework systematically constructs UnrealPose-1M, a dataset comprising approximately one million frames, enriched with multi-view imagery, occlusion annotations, visibility flags, 2D/3D keypoints, and camera parameters. The effectiveness of this synthetic data is validated across four tasks: 3D pose estimation, 2D keypoint detection, 2D-to-3D lifting, and human instance detection and segmentation. The entire dataset and associated tools are publicly released to advance research in label-free or weakly supervised human pose estimation.

Technology Category

Computer Vision: Biometrics, Face, Gesture & PoseNatural Language Processing: Code Generation / Program Synthesis from Natural LanguageHumans and AI: Game Design — Virtual Humans, NPCs and Autonomous Characters

Application Category

Economics, Online Markets and Human Computation: Data quality aspects of human-annotated datasetsWeb Mining and Content Analysis: Web data generation and simulationResponsible Web: Machine-in-the-loop, human agency and autonomy
📝 Abstract
Diverse, accurately labeled 3D human pose data is expensive and studio-bound, while in-the-wild datasets lack known ground truth. We introduce UnrealPose-Gen, an Unreal Engine 5 pipeline built on Movie Render Queue for high-quality offline rendering. Our generated frames include: (i) 3D joints in world and camera coordinates, (ii) 2D projections and COCO-style keypoints with occlusion and joint-visibility flags, (iii) person bounding boxes, and (iv) camera intrinsics and extrinsics. We use UnrealPose-Gen to present UnrealPose-1M, an approximately one million frame corpus comprising eight sequences: five scripted"coherent"sequences spanning five scenes, approximately 40 actions, and five subjects; and three randomized sequences across three scenes, approximately 100 actions, and five subjects, all captured from diverse camera trajectories for broad viewpoint coverage. As a fidelity check, we report real-to-synthetic results on four tasks: image-to-3D pose, 2D keypoint detection, 2D-to-3D lifting, and person detection/segmentation. Though time and resources constrain us from an unlimited dataset, we release the UnrealPose-1M dataset, as well as the UnrealPose-Gen pipeline to support third-party generation of human pose data.
Problem

Research questions and friction points this paper is trying to address.

3D human pose
synthetic data
ground truth
pose estimation
data annotation
Innovation

Methods, ideas, or system contributions that make the work stand out.

synthetic human pose
Unreal Engine 5
3D pose estimation
data generation pipeline
ground truth annotation
🔎 Similar Papers
No similar papers found.