Enhancing Public Speaking Skills in Engineering Students Through AI

📅 2025-11-07
📈 Citations: 0
Influential: 0
📄 PDF

career value

226K/year
🤖 AI Summary
Engineering students often receive insufficient public speaking training, with feedback that is fragmented and non-personalized. Method: This study develops the first multimodal AI evaluation system integrating speech analysis, computer vision, and affective computing. It introduces the novel concept of “expressive consistency,” enabling cross-modal joint modeling of verbal content, acoustic features, facial expressions, body posture, and emotional states to quantitatively assess coordination between linguistic and non-linguistic behaviors. The system incorporates large language models (e.g., Gemini Pro), speech signal processing, 3D pose estimation, and fine-grained emotion recognition for end-to-end automated feedback. Contribution/Results: Preliminary experiments show moderate agreement between AI assessments and expert ratings (Spearman’s ρ = 0.58), with Gemini Pro achieving the highest performance. Students demonstrated significant improvements in speech clarity and professionalism after multiple iterative practice sessions.

Technology Category

Application Category

📝 Abstract
This research-to-practice full paper was inspired by the persistent challenge in effective communication among engineering students. Public speaking is a necessary skill for future engineers as they have to communicate technical knowledge with diverse stakeholders. While universities offer courses or workshops, they are unable to offer sustained and personalized training to students. Providing comprehensive feedback on both verbal and non-verbal aspects of public speaking is time-intensive, making consistent and individualized assessment impractical. This study integrates research on verbal and non-verbal cues in public speaking to develop an AI-driven assessment model for engineering students. Our approach combines speech analysis, computer vision, and sentiment detection into a multi-modal AI system that provides assessment and feedback. The model evaluates (1) verbal communication (pitch, loudness, pacing, intonation), (2) non-verbal communication (facial expressions, gestures, posture), and (3) expressive coherence, a novel integration ensuring alignment between speech and body language. Unlike previous systems that assess these aspects separately, our model fuses multiple modalities to deliver personalized, scalable feedback. Preliminary testing demonstrated that our AI-generated feedback was moderately aligned with expert evaluations. Among the state-of-the-art AI models evaluated, all of which were Large Language Models (LLMs), including Gemini and OpenAI models, Gemini Pro emerged as the best-performing, showing the strongest agreement with human annotators. By eliminating reliance on human evaluators, this AI-driven public speaking trainer enables repeated practice, helping students naturally align their speech with body language and emotion, crucial for impactful and professional communication.
Problem

Research questions and friction points this paper is trying to address.

Engineering students lack sustained personalized public speaking training
Comprehensive verbal and non-verbal feedback is time-intensive for humans
Existing systems assess communication aspects separately without integration
Innovation

Methods, ideas, or system contributions that make the work stand out.

AI-driven multi-modal system for public speaking assessment
Combines speech analysis, computer vision, and sentiment detection
Evaluates verbal, non-verbal communication and expressive coherence
A
Amol Harsh
Computer Science and Artificial Intelligence, Plaksha University, Mohali, India
B
Brainerd Prince
Center for Thinking, Language and Communication, Plaksha University, Mohali, India
S
Siddharth Siddharth
Computer Science and Artificial Intelligence, Plaksha University, Mohali, India
Deepan Muthirayan
Deepan Muthirayan
Assistant Professor, Plaksha University
Control TheoryMachine LearningOptimizationGame Theory
K
Kabir S Bhalla
Computer Science and Artificial Intelligence, Plaksha University, Mohali, India
E
Esraaj Sarkar Gupta
Computer Science and Artificial Intelligence, Plaksha University, Mohali, India
S
Siddharth Sahu
Computer Science and Artificial Intelligence, Plaksha University, Mohali, India