🤖 AI Summary
To address the low image quality, limited resolution, and insufficient class coverage in Indian Sign Language (ISL) image generation, this paper proposes a novel conditional Generative Adversarial Network (cGAN) that integrates Progressive Growing of GANs (ProGAN) with Self-Attention Generative Adversarial Networks (SAGAN), enabling the first controllable synthesis of high-resolution, multi-class ISL images—including letters, digits, and vocabulary. Methodologically, we design a class-conditional attention module and a progressive training strategy to enhance fine-grained detail fidelity and semantic consistency. Our contributions are threefold: (1) construction of the first large-scale, high-quality ISL image dataset comprising 129 vocabulary items; (2) significant performance gains over the ProGAN baseline, achieving +3.2 in Inception Score and −30.12 in Fréchet Inception Distance (FID); and (3) empirical validation that synergistic integration of self-attention and progressive generation substantially improves fine-grained ISL representation learning.
📝 Abstract
Sign language, which contains hand movements, facial expressions and bodily gestures, is a significant medium for communicating with hard-of-hearing people. A well-trained sign language community communicates easily, but those who don't know sign language face significant challenges. Recognition and generation are basic communication methods between hearing and hard-of-hearing individuals. Despite progress in recognition, sign language generation still needs to be explored. The Progressive Growing of Generative Adversarial Network (ProGAN) excels at producing high-quality images, while the Self-Attention Generative Adversarial Network (SAGAN) generates feature-rich images at medium resolutions. Balancing resolution and detail is crucial for sign language image generation. We are developing a Generative Adversarial Network (GAN) variant that combines both models to generate feature-rich, high-resolution, and class-conditional sign language images. Our modified Attention-based model generates high-quality images of Indian Sign Language letters, numbers, and words, outperforming the traditional ProGAN in Inception Score (IS) and Fréchet Inception Distance (FID), with improvements of 3.2 and 30.12, respectively. Additionally, we are publishing a large dataset incorporating high-quality images of Indian Sign Language alphabets, numbers, and 129 words.