🤖 AI Summary
This work addresses the challenge of deploying talking-head generation in resource-constrained educational settings, where existing methods relying on GPUs, large datasets, or complex models are impractical. The authors propose a purely symbolic, CPU-oriented lightweight framework that first converts speech into a phoneme stream and maps it to a compact viseme set. Inspired by the Vedic sutra “Urdhva Tiryakbhyam,” they introduce symbolic coarticulation rules to generate smooth viseme trajectories. Mouth animation is then synthesized via region-of-interest (ROI) deformation and lightweight 2D rendering. This approach, the first to incorporate Vedic computational principles into talking-head generation, achieves high lip-sync accuracy, temporal stability, and identity consistency without deep learning. It operates efficiently on CPU-only systems, significantly reducing computational overhead and latency while outperforming existing CPU-feasible baselines.
📝 Abstract
Talking-head avatars are increasingly adopted in educational technology to deliver content with social presence and improved engagement. However, many recent talking-head generation (THG) methods rely on GPU-centric neural rendering, large training sets, or high-capacity diffusion models, which limits deployment in offline or resource-constrained learning environments. A deterministic and CPU-oriented THG framework is described, termed Symbolic Vedic Computation, that converts speech to a time-aligned phoneme stream, maps phonemes to a compact viseme inventory, and produces smooth viseme trajectories through symbolic coarticulation inspired by Vedic sutra Urdhva Tiryakbhyam. A lightweight 2D renderer performs region-of-interest (ROI) warping and mouth compositing with stabilization to support real-time synthesis on commodity CPUs. Experiments report synchronization accuracy, temporal stability, and identity consistency under CPU-only execution, alongside benchmarking against representative CPU-feasible baselines. Results indicate that acceptable lip-sync quality can be achieved while substantially reducing computational load and latency, supporting practical educational avatars on low-end hardware. GitHub: https://vineetkumarrakesh.github.io/vedicthg