Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of systematic evaluation methods for large language models (LLMs) in educational contexts, which has led to overreliance on subjective judgments by learning engineers. To bridge this gap, the work introduces a trustworthiness framework tailored to education through longitudinal co-design, integrating educational theory with LLM trustworthiness metrics. The authors propose a teaching-oriented trustworthiness assessment framework comprising five dimensions and twenty specific indicators, accompanied by an interactive visualization tool that maps trustworthiness violations to individual model responses. This approach significantly improves inter-rater reliability and enables consistent, comparable evaluations of the pedagogical appropriateness of LLM-generated content, thereby establishing new design principles for education-focused LLM evaluation tools.
📝 Abstract
LLMs are reshaping educational technology, yet evaluating their responses for pedagogical alignment remains underexplored, relying heavily on the expertise of learning engineers building the technology. To bridge this gap, we explore trustworthiness as a structured lens for evaluation, leveraging existing measures of LLM trustworthiness to systematically identify potential pedagogical disruptions. Through a longitudinal co-design process with learning engineers developing an LLM-powered digital textbook, we: (1) co-constructed five trustworthiness metrics comprising 20 measures tailored to pedagogical use; (2) designed visualizations that map trustworthiness violations onto LLM responses; and (3) evaluated how these tools help learning engineers make A/B comparisons of LLM responses. Making trustworthiness explicit increased inter-rater reliability while helping learning engineers resolve conflicting objectives and produce more consistent judgments. We discuss the emergent benefits of trustworthiness as a lens for evaluating LLMs in education and propose new design guidelines for future evaluation tools that enable pedagogically-aligned, LLM-powered learning tools.
Problem

Research questions and friction points this paper is trying to address.

trustworthiness
large language models
educational technology
pedagogical alignment
evaluation metrics
Innovation

Methods, ideas, or system contributions that make the work stand out.

trustworthiness
co-design
LLM evaluation
pedagogical alignment
visualization
A
Adam Coscia
Georgia Institute of Technology, USA
S
Sujata Duwal
Georgia Institute of Technology, USA
L
Langdon Holmes
Vanderbilt University, USA
S
Scott Crossley
Vanderbilt University, USA
Alex Endert
Alex Endert
Associate Professor, Georgia Tech
human-computer interactioninformation visualizationvisual analytics