🤖 AI Summary
To address the high computational complexity, large inference latency, and poor compatibility with multimodal systems exhibited by existing ResNet-based lip-reading models, this paper proposes a lightweight visual speech encoder. Methodologically, it introduces, for the first time, a hierarchical shifted-window mechanism into lip-reading modeling—integrating Swin Transformer architecture, spatiotemporal tokenization, and progressive downsampling to enable joint local–global spatiotemporal feature learning; additionally, a contrastive lip-articulation alignment pretraining strategy is designed to enhance robustness of temporal speech representations. Evaluated on LRW and LRS2 benchmarks, the model achieves state-of-the-art accuracy while reducing parameter count by 37% and accelerating inference by 2.1×, thereby significantly balancing accuracy, efficiency, and deployability.