Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education
2026-08-04 • Human-Computer Interaction
Human-Computer Interaction
AI summaryⓘ
The authors studied how to better evaluate large language models (LLMs) used in educational tools to ensure they teach effectively. They worked with learning engineers to create trustworthiness metrics specifically for education, designed ways to visualize problems in LLM answers, and tested these tools to help engineers compare responses more reliably. Their approach helped make evaluations more consistent and balanced. They suggest using trustworthiness as a useful way to judge LLMs in education and share ideas for future tools to improve this process.
Large Language Models (LLMs)Educational TechnologyTrustworthiness MetricsPedagogical AlignmentLearning EngineersA/B TestingEvaluation ToolsDigital TextbooksInter-rater ReliabilityCo-design
Authors
Adam Coscia, Sujata Duwal, Langdon Holmes, Scott Crossley, Alex Endert
Abstract
LLMs are reshaping educational technology, yet evaluating their responses for pedagogical alignment remains underexplored, relying heavily on the expertise of learning engineers building the technology. To bridge this gap, we explore trustworthiness as a structured lens for evaluation, leveraging existing measures of LLM trustworthiness to systematically identify potential pedagogical disruptions. Through a longitudinal co-design process with learning engineers developing an LLM-powered digital textbook, we: (1) co-constructed five trustworthiness metrics comprising 20 measures tailored to pedagogical use; (2) designed visualizations that map trustworthiness violations onto LLM responses; and (3) evaluated how these tools help learning engineers make A/B comparisons of LLM responses. Making trustworthiness explicit increased inter-rater reliability while helping learning engineers resolve conflicting objectives and produce more consistent judgments. We discuss the emergent benefits of trustworthiness as a lens for evaluating LLMs in education and propose new design guidelines for future evaluation tools that enable pedagogically-aligned, LLM-powered learning tools.