SignGPT: Toward LLM-Mediated Sign Language Interaction through Gloss-Free Translation and Generation

📅 2026-09-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决大语言模型对手语交互支持不足的问题,本文提出SignGTP框架,通过无词汇表的手语翻译与生成统一处理方法来实现双向建模。
📝 Abstract
Large language models (LLMs) provide limited support for sign language interaction. Unifying sign language translation (SLT) and generation (SLG) to enable sign language as both input and output can reduce switching between separate models during sign-text interaction. We present SignGPT, a unified, pose-based framework for gloss-free SLT and SLG. SignGPT integrates part-aware hierarchical representations of body, hand, and facial motion into a shared language model and employs asymmetric multi-token prediction and progressive training for bidirectional modeling. We evaluate SignGPT on How2Sign (ASL) and Phoenix-2014T (DGS) through benchmark comparisons, qualitative analyses, and component ablations. An exploratory study with 12 Deaf ASL signers assesses an LLM-mediated sign-to-sign response pipeline, highlighting the potential of unified modeling to support sign language conversation (SLC).
Problem

Research questions and friction points this paper is trying to address.

sign language interaction
large language models
unified modeling
gloss-free translation and generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

unified framework
gloss-free
asymmetric multi-token prediction
progressive training