🤖 AI Summary
This study addresses the evaluation challenges in humanoid robot sign language translation caused by coupled fitting, repair, and control errors. We construct the HumanoidCSL-20K benchmark, providing four-stage aligned data from source videos to robot trajectories. By introducing a cross-representation tracing mechanism and a kinematic reference protocol, this work presents the first decoupled analysis of motion continuity, feasibility, and execution error. Furthermore, integrating local motion repair, full geometric repair, and multi-granularity scoring enables end-to-end traceable evaluation. Experimental results demonstrate that the proposed approach significantly reduces anomalous gaits and limb penetrations, effectively distinguishes reference learnability from curriculum effects, and precisely identifies execution bottlenecks.
📝 Abstract
Sign data collection is costly, and teleoperation scales poorly, motivating reuse of large video corpora. Humanoid signing requires converting video-derived human motion into robot trajectories while preserving linguistic motion cues. Errors from fitting, human-motion repair, retargeting, robot geometry repair, and control are hard to separate from the final trajectory alone. We introduce HumanoidCSL-20K, a dataset and benchmark of 20,648 sentence-level Chinese Sign Language sequences, each with four aligned versions: the source, the repaired human motion, the direct robot reference, and the geometry-repaired robot reference. Observation-supported local human-motion repair, full-robot geometry repair, and cross-representation provenance make each transformation traceable. Paired evaluations measure human-motion continuity and content preservation, robot-reference feasibility, and physical execution. A sign-specific kinematic-reference protocol scores handshape, location, palm orientation, and inter-hand relation over the full planned motion. Full-corpus results show fewer abnormal arm / hand steps and less inter-hand and hand-body penetration after repair. Control experiments separate reference learnability from curriculum effects, while component scores expose remaining execution errors.