From Sign Language Generation to Humanoid Execution: Vision-Language Guided Retargeting with Collision Mitigation

📅 2026-07-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of infeasible inverse kinematics and unstable trajectories arising from self-collisions when 3D sign language motions generated by data-driven models are executed on humanoid robots. To this end, the authors propose a system-level mapping framework that integrates SMPL-X-based volumetric collision detection with projection-based optimization and task-space retargeting. Notably, they introduce a vision-language model (VLM) as a perceptual critic for the first time, establishing a closed-loop feedback mechanism that jointly enforces physical feasibility and semantic consistency in joint space. Experimental results demonstrate that the proposed approach significantly reduces self-collisions and enhances the executability and stability of sign language gestures on real robotic platforms, thereby validating the critical role of perception-guided control and physical constraints in embodied expression.
📝 Abstract
Recent sign language generation (SLG) systems increasingly output dense 3D body representations, which better preserve full-body kinematics and geometry for downstream embodiment on humanoid robots. However, these generated motions frequently exhibit self-intersections such as hand-hand and hand-torso penetration. While such artifacts may be tolerated in offline rendering, they become critical in humanoid execution as they lead to infeasible inverse-kinematics (IK) solutions, collisions, and unstable retargeted trajectories. We present a system-level framework that bridges SLG outputs to humanoid joint-space execution via two components. First, we introduce a volumetric SMPL-X collision-mitigation module that projects generated signing motions toward physically plausible configurations while minimally deviating from the original trajectory. Second, we propose a vision-language-guided retargeting algorithm built on an IK backbone: a VLM serves as a visual critic over rendered humanoid motion, identifies embodiment-specific failure modes, and triggers targeted task-space corrections. Our results highlight collision handling and perception-guided refinement as key missing components for reliable humanoid signing.
Problem

Research questions and friction points this paper is trying to address.

sign language generation
humanoid execution
collision mitigation
self-intersection
inverse kinematics
Innovation

Methods, ideas, or system contributions that make the work stand out.

collision mitigation
vision-language model
sign language generation
humanoid retargeting
SMPL-X
🔎 Similar Papers
No similar papers found.