RoboSTAR: Next-Scale Autoregressive Sign Language Translation for Humanoid Robots

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scalability limitations of sign language translation and the inability of single temporal granularities to jointly capture long-range and local dependencies. To this end, we propose a text-conditioned, multi-scale autoregressive framework for sign language generation. Methodologically, we introduce partial finite scalar quantization alongside a next-scale prediction mechanism to parallelize the decoding of torso and hand tokens, while enhancing robustness through self-conditioning. Furthermore, by integrating motion retargeting techniques, human movements are transferred to robotic execution, with support for automatic speech recognition (ASR) frontend integration. Experimental results demonstrate that the proposed framework significantly improves the fidelity of generated motions, effectively enhancing the accessibility interaction capabilities of humanoid robots in public scenarios.
📝 Abstract
Sign-language interpretation in public communication relies on qualified professional interpreters and can be difficult to scale, motivating robotic signing as a complementary accessibility interface. We present RoBoSTAR, a text-conditioned sign language production (SLP) framework for generating human-centric sign motion that can be retargeted for robotic execution, with speech supported optionally through an external ASR front end. Conventional autoregressive approaches flatten motion into a single full-resolution token sequence, forcing long-range and local dependencies to be modeled at a uniform temporal granularity. RoBoSTAR instead combines part-wise Finite Scalar Quantization with next-scale autoregression, generating motion over progressively finer temporal resolutions while predicting synchronized body and hand tokens in parallel within each step. This coarse-to-fine formulation provides compact long-range context before progressively refining motion details, while self-conditioning and context corruption improve robustness to cross-scale prediction errors. The generated motion is subsequently retargeted for physical humanoid execution. Extensive qualitative and quantitative evaluations are conducted to demonstrate the effectiveness of RoBoSTAR.
Problem

Research questions and friction points this paper is trying to address.

Sign Language Production
Humanoid Robots
Autoregressive Modeling
Motion Generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Next-Scale Autoregression
Finite Scalar Quantization
Sign Language Production
Coarse-to-Fine Generation
Motion Retargeting
💼 Related Jobs
No related jobs found.