Capabilities of Claude Fable 5 on Biomedical Challenge Problems
2026-07-12 • Computation and Language
Computation and Language
AI summaryⓘ
The authors tested a new AI language model, Claude Fable 5, on several biomedical question sets to see how well it performs. They found that this model often refuses to answer many questions, especially in basic science and rare disease topics, unlike older models and GPT-5. When it does answer, Claude Fable 5 is as accurate or better than the other models. So, the main issue with this model is that it often chooses not to respond, rather than giving wrong answers.
language modelbiomedical benchmarksClaude Fable 5answer refusalMedQAMedXpertQARareBenchdeterministic scoringGPT-5
Authors
Dominic Okonkwo, Magnus Hodgson, Temitope I. David, Susan Adanna Ihejirika
Abstract
Frontier language models are increasingly evaluated on biomedical benchmarks, but two problems undermine most published evaluations: legacy benchmarks are near-saturated, and open-ended responses are graded by other language models. We evaluate Claude Fable 5, Anthropic's most capable publicly available model, across eight biomedical benchmarks, four text and four multimodal, using deterministic scoring against fixed answer keys throughout. We include two Claude predecessors and GPT-5 as baselines. Refusal is tracked as a distinct outcome in every result table. That decision produces the paper's central finding. Fable 5 refuses between 8.0% and 99.4% of questions depending on the benchmark, a pattern absent in both predecessors and in GPT-5. Once refused items are excluded from the denominator, Fable 5's accuracy exceeds or meets every other model on every benchmark in this study. We identify two distinguishable refusal patterns: one concentrating in basic-science and mechanism content across MedQA and MedXpertQA MM, confirmed independently on two benchmarks using each benchmark's own category labels; and a separate disease-domain pattern on RareBench, where inborn metabolic disease presentations are refused near-universally while adult-onset autoimmune presentations are not. The primary constraint on Fable 5's biomedical usefulness is willingness to engage, not capability once it does.