The complexities of patient-centred conversational artificial intelligence
2026-07-09 • Artificial Intelligence
Artificial IntelligenceComputation and Language
AI summaryⓘ
The authors studied over 2,000 real conversations between patients and health chatbots and found people communicate very differently, often showing a wide range of emotions. They created a new patient simulator that mimics clinical details, emotions, and communication styles separately. In tests, humans had a hard time telling simulated conversations from real ones. When testing chatbots’ ability to judge how urgent a health issue was, they discovered that how a patient talks can change the chatbot’s decision. The authors suggest that health chatbots need to handle many communication styles to work well for everyone and avoid unfair results.
Large Language ModelsHealth ChatbotsSymptom AssessmentPatient SimulatorCommunication StyleEmotional ExpressionTriageUrgency AssessmentTuring TestHealthcare Disparities
Authors
João Matos, Olivia Buege, Donny Cheung, Gary S. Collins, Paula Dhiman, Nan Li, Bingyu Mao, Benjamin W. Nelson, Michail Ouroutzoglou, Paul Varghese, Jonathan Amar
Abstract
Consumer-facing health chatbots powered by large language models (LLMs) are increasingly used for symptom assessment. However, chatbot development and evaluation often rely on cooperative, articulate, simulated patients. We analysed 2,053 real patient-chatbot conversations and found that communication patterns and expression of emotions vary widely across users. We developed a patient simulator that separately models clinical content, emotional state, conversational strategy, and communication style. In a Turing-inspired evaluation of realism with 15 human graders, simulated conversations were nearly indistinguishable from real ones, with human graders achieving an accuracy of 55%. We used five distinct patient personae, across 1,164 clinician-graded cases, to evaluate the performance of four LLMs in urgency assessment. We found that communication style can significantly alter triage outcomes. Patient-centred conversational artificial intelligence must accommodate communication diversity: systems designed for idealised, rather than realistic, interactions risk underperforming and amplifying health disparities when deployed in the real world.