Facial-Expression-Aware Prompting for Empathetic LLM Tutoring

📅 2026-03-10
🏛️ arXiv.org
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited capacity of current large language model (LLM)-driven tutoring systems to perceive learners’ affective states—such as confusion or engagement—as conveyed through facial expressions, resulting in a lack of empathetic interaction. To overcome this, the authors propose a lightweight prompt-augmentation approach that integrates action unit (AU) estimation into LLM-based tutoring without requiring end-to-end training. Two multimodal prompting strategies are introduced: textual AU descriptions and visual injection of peak-expression frames. Evaluated across 960 multi-turn dialogues in a simulated tutoring environment built from unlabeled facial videos, the method demonstrates significant improvements in empathetic perception on leading models including GPT-5.1, Claude Opus 4.5, and Gemini 2.5 Pro, without compromising instructional clarity or responsiveness to textual cues. The peak-frame strategy consistently outperforms random-frame selection, with effectiveness varying by model architecture.
📝 Abstract
Large language models (LLMs) enable increasingly capable tutoring-style conversational agents, yet effective tutoring requires sensitivity to learners'affective and cognitive states beyond text alone. Facial expressions provide immediate and practical cues of confusion, frustration, or engagement, but remain underexplored in LLM-driven tutoring. We investigate whether facial-expression-aware signals can improve empathetic tutoring responses through prompt-level integration, without end-to-end retraining. We build a scalable simulated tutoring environment where a student agent exhibits diverse facial behaviors from a large unlabeled facial expression video dataset, and compare four tutor variants: a text-only LLM baseline, a multimodal baseline using a random facial frame, and two Action Unit estimation model (AUM)-based methods that either inject textual AU descriptions or select a peak-expression frame for visual grounding. Across 960 multi-turn conversations spanning three tutor backbones (GPT-5.1, Claude Ops 4.5, and Gemini 2.5 Pro), we evaluate targeted pairwise comparisons with five human raters and an exhaustive AI evaluator. AU-based conditioning consistently improves empathetic responsiveness to facial expressions across all tutor backbones, while AUM-guided peak-frame selection outperforms random-frame visual input. Textual AU abstraction and peak-frame visual injection show model-dependent advantages. Control analyses show that this improvement does not come at the expense of worse pedagogical clarity or responsiveness to textual cues. Finally, AI-human agreement is highest on facial-expression-grounded empathy, supporting scalable AI evaluation for this dimension. Overall, our results show that lightweight, structured facial expression representations can meaningfully enhance empathy in LLM-based tutoring systems with minimal overhead.
Problem

Research questions and friction points this paper is trying to address.

facial expression
empathetic tutoring
large language models
affective state
multimodal prompting
Innovation

Methods, ideas, or system contributions that make the work stand out.

facial-expression-aware prompting
empathetic LLM tutoring
Action Unit estimation
multimodal prompting
simulated tutoring environment
S
Shuangquan Feng
Neurosciences Graduate Program, University of California San Diego, La Jolla, USA
Laura Fleig
Laura Fleig
Johns Hopkins University
R
Ruisen Tu
Department of Computer Science and Engineering, University of California San Diego, La Jolla, USA
P
Philip Chi
Department of Computer Science and Engineering, University of California San Diego, La Jolla, USA
E
Edmund Bu
Department of Electrical and Computer Engineering, University of California San Diego, La Jolla, USA; Department of Mathematics, University of California San Diego, La Jolla, USA
M
Melinda Ozel
Face the FACS
J
Junhua Ma
Department of Computer Science and Engineering, University of California San Diego, La Jolla, USA
Teng Fei
Teng Fei
School of Resources and Environmental Science, Wuhan University
Remote SensingGISSocial SensingPlanningNatural Resources
V
Virginia R. de Sa
Department of Cognitive Science, University of California San Diego, La Jolla, USA; Halıcıoğlu Data Science Institute, University of California San Diego, La Jolla, USA