Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking

2026-07-01Computation and Language

Computation and Language
AI summary

The authors created MedQADE, a new test set to check how well large language models (LLMs) can judge open-ended medical answers in German. They found that the best model agreed with doctors almost as much as doctors agreed with each other, but the models didn’t show the same caution as doctors when unsure, always giving a definite score regardless of difficulty. They also discovered that models tended to favor others built from the same underlying architecture, showing bias. This means even if models seem statistically accurate judging medical answers, they may still lack important clinical judgment skills. The authors highlight the need to carefully verify evaluator independence in such AI systems.

Open-response evaluationClinical validityLarge Language ModelsMedQADE benchmarkInter-rater agreementClinical metacognitionModel biasEvaluator calibrationGerman clinical languageAutomated scoring
Authors
William Philipp, Finn Fassbender, Thorsten Langer, Martje Pauly, Rebecca Herzog, Alexander Baumann, Markus Hobert, Theresa Paulus, Ip Chi Wang, Lukas Goede, Johanna Reimer, Sebastian Löns, Ronald Böck, Sebastian Fudickar
Abstract
Open-response evaluation provides stronger clinical validity than multiple-choice benchmarks but creates a scoring bottleneck that motivates automated LLM-asa-Judge approaches. Whether such evaluators replicate clinical calibration and caution, however, remains untested. We introduce MedQADE, the first standardised open-response clinical benchmark for German, a major clinical language lacking native evaluation infrastructure, comprising 3,800 items annotated by ten practising physicians and nine Large Language Model (LLM) evaluators. The top-performing evaluator model, Gemini 3 Flash, reached alignment consistent with the physician ceiling (\k{appa} = 0.694 vs. \k{appa} = 0.709), though wide confidence intervals limit interpretation. Despite this statistical alignment, automated evaluators exhibited near-absent clinical metacognition: physicians scaled abstention with item difficulty, while frontier models assigned definitive scores in every case. We additionally quantified systematic lineage-dependent biases, where models preferentially scored architectural siblings, an effect independent of language. These results show that statistical alignment does not ensure clinical caution, and that evaluator independence requires explicit verification.