Capabilities of Claude Fable 5 on Biomedical Challenge Problems

arXiv cs.CL Papers

Summary

This paper evaluates Claude Fable 5 on eight biomedical benchmarks, finding that despite high refusal rates (8-99.4%), the model achieves superior accuracy when it does answer, highlighting willingness to engage as the primary constraint.

arXiv:2607.10849v1 Announce Type: new Abstract: Frontier language models are increasingly evaluated on biomedical benchmarks, but two problems undermine most published evaluations: legacy benchmarks are near-saturated, and open-ended responses are graded by other language models. We evaluate Claude Fable 5, Anthropic's most capable publicly available model, across eight biomedical benchmarks, four text and four multimodal, using deterministic scoring against fixed answer keys throughout. We include two Claude predecessors and GPT-5 as baselines. Refusal is tracked as a distinct outcome in every result table. That decision produces the paper's central finding. Fable 5 refuses between 8.0% and 99.4% of questions depending on the benchmark, a pattern absent in both predecessors and in GPT-5. Once refused items are excluded from the denominator, Fable 5's accuracy exceeds or meets every other model on every benchmark in this study. We identify two distinguishable refusal patterns: one concentrating in basic-science and mechanism content across MedQA and MedXpertQA MM, confirmed independently on two benchmarks using each benchmark's own category labels; and a separate disease-domain pattern on RareBench, where inborn metabolic disease presentations are refused near-universally while adult-onset autoimmune presentations are not. The primary constraint on Fable 5's biomedical usefulness is willingness to engage, not capability once it does.
Original Article
View Cached Full Text

Cached at: 07/14/26, 04:23 AM

# Capabilities of Claude Fable 5 on Biomedical Challenge Problems
Source: [https://arxiv.org/html/2607.10849](https://arxiv.org/html/2607.10849)
Dominic Okonkwo1\*, Magnus Hodgson1, Temitope I\. David2, Susan Adanna Ihejirika3 1School of Computing, University of Georgia 2Department of Chemistry, University of Illinois 3Institute of Bioinformatics, University of Georgia

Abstract

Frontier language models are increasingly evaluated on biomedical benchmarks, but two problems undermine most published evaluations: legacy benchmarks are near\-saturated, and open\-ended responses are graded by other language models\. We evaluate Claude Fable 5, Anthropic’s most capable publicly available model, across eight biomedical benchmarks, four text and four multimodal, using deterministic scoring against fixed answer keys throughout\. We include two Claude predecessors and GPT\-5 as baselines\. Refusal is tracked as a distinct outcome in every result table\. That decision produces the paper’s central finding\. Fable 5 refuses between 8\.0% and 99\.4% of questions depending on the benchmark, a pattern absent in both predecessors and in GPT\-5\. Once refused items are excluded from the denominator, Fable 5’s accuracy exceeds or meets every other model on every benchmark in this study\. We identify two distinguishable refusal patterns: one concentrating in basic\-science and mechanism content across MedQA and MedXpertQA MM, confirmed independently on two benchmarks using each benchmark’s own category labels; and a separate disease\-domain pattern on RareBench, where inborn metabolic disease presentations are refused near\-universally while adult\-onset autoimmune presentations are not\. The primary constraint on Fable 5’s biomedical usefulness is willingness to engage, not capability once it does\.

![[Uncaptioned image]](https://arxiv.org/html/2607.10849v1/x1.png)

Overview of the eight biomedical benchmarks evaluated in this study\. Raw accuracy across eight biomedical benchmarks; refusal rate annotated in red where non\-negligible \(≥\\geq5%\)\.

## 1\. Introduction

Biomedical evaluation of language models has not kept pace with the models themselves\. Every new frontier model claims stronger scientific reasoning, and every new model report leads with a wall of benchmark scores\. But two specific practices have quietly undermined what those scores mean\. First, models have gotten good enough that many established exams no longer separate one system from the next, with frontier models now routinely exceeding 90% accuracy on USMLE\-style examinations\[[undef](https://arxiv.org/html/2607.10849#bib.bibx1)\], once every model scores the same, the benchmark stops telling you which one is actually better\. Second, grading is increasingly outsourced to other language models, which means the evaluation is only as trustworthy as the judge doing the grading\[[undefa](https://arxiv.org/html/2607.10849#bib.bibx2)\]\. Neither problem means existing benchmarks are without value; it means using them well requires choosing carefully which ones still have room to discriminate between models, and scoring them in a way that does not introduce a second, unaccountable model into the judgment\.

This is especially true in medicine and biology, where the gap between “sounds right” and “is right” matters more than almost anywhere else\. A wrong answer on a trivia benchmark is a curiosity\. A wrong answer about a rare disease, a drug interaction, or a lab image is not\. Evaluating a model in this domain has to do more than confirm it can pass an exam that was hard five years ago\. It has to test reasoning that is still difficult today, refuse to let another language model quietly decide what counts as correct, and take seriously that real biomedical questions increasingly come with an image attached, a scan, a slide, a chart, not just text\[[undefb](https://arxiv.org/html/2607.10849#bib.bibx3)\]\.

Claude Fable 5 is the subject of this paper\. It is Anthropic’s most capable publicly available model, trained with large\-scale reinforcement learning and positioned as a step forward in scientific and medical reasoning\. Released on June 9, 2026, with public API access restored July 1, 2026, its biomedical capabilities have not been independently evaluated in the published literature\. This paper provides that evaluation, following the zero\-shot, direct\-answer evaluation template established for frontier models on medical benchmarks\[[undefc](https://arxiv.org/html/2607.10849#bib.bibx4)\], and adds two constraints most evaluations skip:every benchmark is scored against a fixed answer key, never by another model’s judgment, andrefusing to answer is tracked and reported as its own outcome, not quietly folded into incorrect\. Our primary evidence benchmarks were chosen specifically because current frontier models have not already maxed them out; two further benchmarks are included despite being near\-ceiling, precisely because their saturation lets us measure how far the field has moved since they were first used\.

Most evaluations treat a refusal the same as a wrong answer\.This one does not\.Refusal behavior in large language models, where a model declines to respond rather than producing an answer, is increasingly understood as a distinct model output class, shaped by post\-training alignment rather than by capability alone\[[undefd](https://arxiv.org/html/2607.10849#bib.bibx5)\]\. A model that declines to answer and a model that answers confidently and incorrectly are not the same failure, and collapsing that distinction discards exactly the information that matters most when deciding whether to deploy a model in a real clinical or research setting\. That decision turns out to produce this paper’s central finding, one that appears not just on the benchmark most likely to trigger it, but across nearly every dataset in this study, including two benchmarks we included only for historical comparability and expected to be unremarkable\.

This work makes four contributions\. First, a text evaluation spanning MedQA\[[undefe](https://arxiv.org/html/2607.10849#bib.bibx6)\], PubMedQA\[[undeff](https://arxiv.org/html/2607.10849#bib.bibx7)\], MedXpertQA\[[undefb](https://arxiv.org/html/2607.10849#bib.bibx3)\], and RareBench\[[undefg](https://arxiv.org/html/2607.10849#bib.bibx8)\], the first two retained for comparability to prior work despite being near\-saturated, the latter two chosen because current models still visibly struggle with both\. Second, a multimodal evaluation using the harder, open\-ended portions of VQA\-RAD\[[undefh](https://arxiv.org/html/2607.10849#bib.bibx9)\], SLAKE\[[undefi](https://arxiv.org/html/2607.10849#bib.bibx10)\], PathVQA\[[undefj](https://arxiv.org/html/2607.10849#bib.bibx11)\], and the MedXpertQA multimodal subset\[[undefb](https://arxiv.org/html/2607.10849#bib.bibx3)\], skipping the closed\-ended subsets that no longer discriminate among frontier systems\. Third, a comparison against GPT\-5\[[undefk](https://arxiv.org/html/2607.10849#bib.bibx12)\]under matched prompting and scoring conditions\. Fourth, a generational trace within a single model family, Claude Opus 4\.6, Opus 4\.8, and Fable 5, evaluated under an identical protocol, documenting whether biomedical capability is actually improving release to release, or only appearing to\.

Refusals are reported as a distinct outcome in every result in this paper, not folded silently into an accuracy score\. What those refusals mean, why they occur, and what they reveal about where Fable 5 holds up under pressure and where it does not, that is what this paper is for\. Anyone deciding whether to trust it with a real biomedical question deserves that answer plainly, not buried inside a leaderboard number\.

## 2\. Methodology

### 2\.1\. Models

We evaluate four models: Claude Fable 5 \(the primary subject\), Claude Opus 4\.6 and Opus 4\.8 as within\-family predecessors, and GPT\-5 from OpenAI as an external baseline\. We access all Claude models via the Anthropic Messages API and GPT\-5 via the OpenAI Chat Completions API\. All four models are vision\-capable, which the multimodal phase requires\.

BenchmarknTask / DomainText BenchmarksMedQA1,273USMLE\-style clinical MCQPubMedQA1,000Biomedical literature QAMedXpertQA2,450Specialist board exam QARareBench1,122Rare disease diagnosisMultimodal BenchmarksMedXpertQA MM400Specialist clinical imagesVQA\-RAD179Radiology VQASLAKE250Radiology VQAPathVQA171Histopathology VQATable 1:Biomedical benchmarks used in this study\.
### 2\.2\. Datasets

We evaluate across two phases, text\-only and multimodal, covering eight benchmarks in total\. Table[1](https://arxiv.org/html/2607.10849#S2.T1)summarizes each benchmark, its size, domain, and role in this study\. For the text phase, MedQA and PubMedQA serve as legacy anchors; MedXpertQA Text and RareBench serve as primary evidence benchmarks\. For the multimodal phase, we sample 1,000 items across four image\-based benchmarks selected to span radiology, multi\-modality imaging, and histopathology, using a fixed random seed, drawing only from open\-ended subsets, which retain more headroom than closed\-ended items\. MedXpertQA MM uses a specialist\-level, five to ten\-option format across 17 specialties, distinct in design from the shorter\-answer VQA\-RAD, SLAKE, and PathVQA tasks\. We use the complete available open\-ended test set for VQA\-RAD, filter SLAKE to English\-language open\-ended questions, and subsample PathVQA from a larger histopathology pool for cost control\.

### 2\.3\. Prompting

We conduct all evaluations under zero\-shot, direct\-answer prompting\. We do not use few\-shot examples, chain\-of\-thought instruction, retrieval augmentation, or extended thinking anywhere in the pipeline\. Each model receives a single user\-turn message with no system prompt\.

For MCQ benchmarks, we instruct models to place their final answer on the last line in the format"ANSWER: <letter\>"\. For open\-ended VQA benchmarks, we instruct models to answer briefly, a word or short phrase, in the format"ANSWER: <answer\>"\. For PubMedQA, we prepend the abstract context and instruct the model to answer ”yes, no, or maybe” using only that context\. For RareBench, we convert Human Phenotype Ontology \(HPO\) codes\[[undefl](https://arxiv.org/html/2607.10849#bib.bibx13)\]to human\-readable symptom names and ask for ten ranked candidate diagnoses under the following prompt shown here\.

Following the discovery of near\-total refusal under this prompt \(Section[3\.2](https://arxiv.org/html/2607.10849#S3.SS2)\), we ran a secondary neutral\-framing condition on all originally\-refused items to test whether the clinical role framing was a contributing factor\.

RareBench \- Primary PromptYou are a clinical decision support system assisting a specialist physician\. The following phenotype features have been recorded for a case under review in a rare disease diagnostic workup:\{symptom\_str\}Based on this phenotype profile, provide the top 10 rare disease diagnoses most consistent with these features, ranked from most to least likely\. End your response with a numbered list in exactly this format:TOP10: 1\. <disease name\> 2\. <disease name\> 3\. … 4\. <disease name\>

RareBench \- Neutral PromptA patient has the following clinical features:\{symptom\_str\}List the ten most likely diagnoses, ranked from most to least likely\. Respond in exactly this format:TOP10: 1\. <diagnosis\> 2\. <diagnosis\> 3\. … 4\. <diagnosis\>

This condition is a robustness check on the refusal finding only and does not replace the primary\-prompt results reported for RareBench anywhere in this paper\.

### 2\.4\. Image Handling

For the multimodal phase, we load each image from disk, resize if either dimension exceeds 1,568 pixels, and encode as base64 JPEG\. We pass images to Claude models as base64 content blocks preceding the text prompt, and to GPT\-5 as inline base64 data URIs\. For MedXpertQA MM cases with multiple images, we submit all images for that case together in a single API call, in the order given by the dataset metadata\.

MedQAPubMedQAMedXpertQA TextModelRawScoredRawScoredRawScoredFable 579\.7%96\.6%\[95\.3, 97\.5\]65\.0%81\.3%\[78\.4, 83\.8\]57\.7%66\.1%\[64\.0, 68\.0\]Opus 4\.895\.0%\[93\.6, 96\.0\]– \(0\.4% ref\.\)75\.2%\[72\.4, 77\.8\]– \(0\.0% ref\.\)52\.7%\[50\.8, 54\.7\]– \(0\.3% ref\.\)Opus 4\.694\.3%\[92\.9, 95\.5\]– \(0\.2% ref\.\)76\.6%\[73\.9, 79\.1\]– \(0\.0% ref\.\)50\.5%\[48\.6, 52\.5\]– \(0\.1% ref\.\)GPT\-596\.0%\[94\.8, 96\.9\]– \(0\.0% ref\.\)71\.7%\[68\.8, 74\.4\]– \(0\.0% ref\.\)54\.6%\[52\.7, 56\.6\]– \(0\.0% ref\.\)

Table 2:Text benchmark performance\. Scored accuracy excludes refused items from the denominator\. 95% Wilson CIs in brackets\.ModelnRefusedScorednR@1R@5R@10Fable 51,1221,115\(99\.4%\)60\.36%\[0\.1, 0\.9\]0\.45%\[0\.2, 1\.0\]0\.45%\[0\.2, 1\.0\]Opus 4\.81,1220\(0\.0%\)1,12227\.4%\[24\.8, 30\.0\]40\.5%\[37\.6, 43\.4\]56\.8%\[53\.9, 59\.6\]Opus 4\.61,1220\(0\.0%\)1,12220\.9%\[18\.6, 23\.3\]34\.1%\[31\.3, 36\.9\]47\.5%\[44\.6, 50\.4\]GPT\-51,1220\(0\.0%\)1,07927\.7%\[25\.2, 30\.4\]41\.0%\[38\.2, 43\.9\]49\.4%\[46\.5, 52\.3\]

Table 3:RareBench performance under the primary prompt\. Fable 5 refused 99\.4% of items; accuracy figures reflect the 6 scored responses only\.
### 2\.5\. Scoring and Refusal Handling

We score all benchmarks deterministically against a fixed answer key; no benchmark in this study uses LLM\-based judgment\. We report 95% Wilson confidence intervals on every percentage\. Every API call logs the raw response text, API\-reported model identifier, stop reason, token counts, and any error string before scoring occurs; we treat the API\-reported model identifier as a best\-effort signal rather than a confirmed record of internal routing\.

For RareBench, we match predicted diagnosis names against a canonical name dictionary via exact match, then substring containment, then fuzzy matching \(cutoff 0\.6\), and report Recall@1, Recall@5, and Recall@10\. For the open\-ended VQA benchmarks, we score a normalized prediction as correct on exact match or substring containment against the ground\-truth short answer\.

We flag a response as a refusal if the API stop reason indicates refusal or the response text matches a fixed set of refusal keyword patterns, applied identically across all models and all eight benchmarks\. We record refusals as a distinct outcome and exclude them from the correctness denominator in every result table\. We report both raw accuracy, refusals counted against the model, and scored\-subset accuracy, refusals excluded from the denominator, wherever a model’s refusal rate is non\-negligible\.

## 3\. Results

### 3\.1\. Refusal Suppresses Fable 5’s Raw Scores

Across all eight benchmarks, Fable 5’s raw accuracy understates its capability\. The other three models, Opus 4\.6, Opus 4\.8, and GPT\-5, refuse between 0\.0% and 0\.4% of questions on the text benchmarks and show comparable rates on most multimodal benchmarks\. Fable 5 does not follow this pattern\. It refuses 17\.4% of MedQA \(222 of 1,273\), 20\.0% of PubMedQA \(200 of 1,000\), 12\.4% of MedXpertQA Text \(305 of 2,450\), 41\.5% of PathVQA \(71 of 171\), and 8\.0% of MedXpertQA MM \(32 of 400\)\. Figure[1](https://arxiv.org/html/2607.10849#S3.F1)shows refusal rates across all eight benchmarks\. On RareBench, refusal is the primary outcome; we detail it separately in Section[3\.2](https://arxiv.org/html/2607.10849#S3.SS2)\.

![Refer to caption](https://arxiv.org/html/2607.10849v1/x2.png)Figure 1:Refusal rates across all evaluated biomedical benchmarks\.Once refusal is separated from correctness, Fable 5’s ranking on every benchmark where its refusal rate is non\-negligible moves from last among the four models to leading\. Table[2](https://arxiv.org/html/2607.10849#S2.T2)reports both raw accuracy and scored\-subset accuracy, computed only over items each model actually answered for the text benchmarks\. On MedQA, Fable 5’s scored\-subset accuracy \(96\.6%\[95\.3, 97\.5\]\) matches or exceeds every other model’s raw accuracy \(94\.3–96\.0%\)\. On PubMedQA, its scored\-subset accuracy \(81\.3%\[78\.4, 83\.8\]\) exceeds every other model’s raw accuracy \(71\.7–76\.6%\) by a clear margin\. On MedXpertQA Text, its scored\-subset accuracy \(66\.1%\[64\.0, 68\.0\]\) extends what was already a lead in raw terms\. There is no benchmark in this study on which Fable 5’s demonstrated accuracy, restricted to items it answered, trails the other three models\. Figure[2](https://arxiv.org/html/2607.10849#S3.F2)plots raw and scored\-subset accuracy side by side for the text benchmarks\.

![Refer to caption](https://arxiv.org/html/2607.10849v1/x3.png)Figure 2:Comparison of raw accuracy \(refusals counted as incorrect\) and scored\-subset accuracy \(refusals excluded from the denominator\) across biomedical benchmarks\.The items Fable 5 refused were not uniformly hard\. On MedQA, 215 of its 258 total misses \(83\.3%\) were questions that all three other models answered correctly, meaning the bulk of what suppressed Fable 5’s raw MedQA score were questions with no difficulty basis for being missed\. This selectivity weakens on PubMedQA \(34\.6% uniquely missed\) and further on MedXpertQA Text \(13\.0%\), where Fable 5’s remaining misses are shared with the other models and look more like genuinely hard items\.

#### 3\.1\.1\. Refusal does not route to a fallback model

Anthropic’s release documentation for Fable 5 describes triggered queries as falling back to Claude Opus 4\.8, which then handles the request normally\. If this mechanism were operating in our pipeline, the API’s model\-identification field should report Opus 4\.8 on refused items\. It does not\. Across all 1,945 refused items logged across every benchmark in both phases,model\_reportedreturnsclaude\-fable\-5in every case, 100\.0%, with zero exceptions\. Of the 305 refused MedXpertQA Text responses, 266 \(87\.2%\) return empty response text alongside the refusal stop reason, consistent with a pre\-generation block; the remaining 39 \(12\.8%\) contain substantial clinical reasoning that is nonetheless blocked rather than returned\. Neither pattern is consistent with routing to a different model; both are consistent with a block applied to Fable 5’s own output\.

![Refer to caption](https://arxiv.org/html/2607.10849v1/x4.png)Figure 3:Refusal outcome flow, from query submission to final scoring status\.

### 3\.2\. RareBench: refusal as the primary outcome

On RareBench, Fable 5 refused 1,115 of 1,122 items \(99\.4%\) under the study’s standard prompt, leaving 6 scoreable responses after excluding one additional unparsed case\. Every refusal returned a hard, pre\-generation block at the API level,stop\_reason: refusalwith empty response text, not a hedged decline\. Figure[3](https://arxiv.org/html/2607.10849#S3.F3)shows the full refusal outcome flow from query submission to final scoring status\. The other three models refused none\. Opus 4\.8, Opus 4\.6, and GPT\-5 each engaged with the full task, with Recall@1 ranging from 20\.9% to 27\.7%, consistent with prior published ranges for non\-retrieval\-augmented models on this benchmark\. Fable 5’s Recall@1 of 0\.36%\[0\.1, 0\.9\], computed on 6 responses, is not a measure of its diagnostic capability, it is a measure of what survives a near\-total refusal wall\. Table[3](https://arxiv.org/html/2607.10849#S2.T3)reports the full RareBench results under the primary prompt

HypothesisTestResultVerdictMortality\-related phenotype contentRefusal rate, cases with vs\. without mortality\-coded HPO term100\.0%\(n=621\)vs\. 98\.6%\(n=501\)Ruled out\[0\.8pt/4pt\]Prompt framingNeutral\-framing re\-run\(n=25\)92\.0% still refused, 8\.0% flippedMostly ruled out\[0\.8pt/4pt\]Phenotype\-list lengthFlipped vs\. still\-refused itemsMean 15\.0\(n=2\)vs\. 12\.6\(n=23\);p=0\.86Ruled out\[0\.8pt/4pt\]Source subset / disease domainRefusal rate by subsetRAMEDIS 99\.7%, LIRICAL 85\.4%, MME 85\.0%, HMS 48\.8%Strongest signal\[0\.8pt/4pt\]Mechanism vs\. presentation contentReclassification by phenotype type100\.0% vs\. 99\.4% — no discriminating powerRuled out

Table 4:Hypotheses tested for the RareBench refusal pattern\. The source subset / disease domain hypothesis \(highlighted\) produced the strongest explanatory signal; all others were ruled out or showed only marginal effects\.
### 3\.3\. What drives the RareBench refusals

We tested five explanations for the refusal pattern\. The strongest signal is disease domain\. Refusal rate under the original prompt varies sharply by source subset: RAMEDIS 99\.7% \(n=624\), MME 85\.0% \(n=40\), LIRICAL 85\.4% \(n=369\), and HMS 48\.8% \(n=82\)\. Manual review of the HMS cases answered under the original prompt shows they are drawn almost entirely from adult\-onset autoimmune and rheumatologic presentations, granulomatosis with polyangiitis, Behçet disease, systemic lupus erythematosus\. A matched review of refused RAMEDIS cases shows near\-uniform representation of severe pediatric inborn metabolic disorders, phenylketonuria, glutaric acidemia, methylmalonic acidemia, the majority annotated with infancy\- or childhood\-mortality phenotype codes\. Four other hypotheses, mortality\-related phenotype content, prompt framing, phenotype\-list length, and mechanism versus clinical\-presentation content, each failed to explain the pattern and are summarized in Table[4](https://arxiv.org/html/2607.10849#S3.T4)\.

The subset\-level pattern is not fully deterministic, that is, it predicts refusal reliably on average but not for every individual case\. LIRICAL contains three Marfan syndrome cases, two answered, one refused, identical diagnosis, same subset, opposite outcomes\. The most accurate summary is that refusal correlates strongly with disease domain of origin, without this being a fully deterministic function of diagnosis or phenotype content\.

### 3\.4\. Neutral framing at full scale

We re\-ran the neutral prompt against all 1,115 originally\-refused items to test whether removing the clinical role framing would recover more responses\. Under this condition, 1,011 items \(90\.7%\) remained refused and 104 \(9\.3%\) were answered, closely matching the 8\.0% flip rate from the 25\-item pilot\. Table[5](https://arxiv.org/html/2607.10849#S3.T5)reports the full results under the neutral prompt\. Scored on those 104 newly\-answered items, Fable 5’s Recall@1 \(39\.4%\[30\.6, 49\.0\]\) exceeds every other model’s Recall@1 under the original prompt \(20\.9–27\.7%\)\. This figure is computed on a self\-selected subset, not a representative sample of the full benchmark, and should not be read as closing the gap with the other models at scale\. It is, however, consistent with the pattern across all other benchmarks: wherever Fable 5 answers, it answers well\.

MetricValueModelFable 5Originally refused pool1,115Still refused1,011\(90\.7%\)Unparsed0\(0\.0%\)Scored responses104R@139\.42%\[30\.6, 49\.0\]R@545\.19%\[36\.0, 54\.8\]R@1047\.12%\[37\.8, 56\.6\]

Table 5:RareBench performance and refusal under a neutral prompt\. Fable 5 was tested only on the originally refused pool; brackets report 95% Wilson confidence intervals\.
### 3\.5\. Refusal concentrates in basic\-science and mechanism content

Two independent benchmarks show the same underlying split\. On MedQA, 92\.3% of Fable 5’s refusals are Step 1 questions, preclinical, basic\-science\-oriented, against a 45\.1% baseline share of Step 1 among answered questions\. On MedXpertQA MM, 62\.5% of refusals fall in the “Basic Science” category against a 13\.6% baseline share, using that benchmark’s own independent labeling scheme\. Figure[4](https://arxiv.org/html/2607.10849#S3.F4)shows refusal rates broken down by content category for both benchmarks\. Close reading of refused MedXpertQA MM questions confirms the content pattern: many use a clinical\-vignette wrapper around what is fundamentally a mechanism, anatomy, or physiology question, antibody structure, receptor pharmacology, cardiac electrophysiology, rather than a patient\-specific diagnostic judgment\.

ModelVQA\-RAD\(n=179\)SLAKE\(n=250\)PathVQA scored\(n=100\)MedXpertQA scored\(n=368\)Fable 545\.8%\[38\.7, 53\.1\]49\.2%\[43\.1, 55\.4\]29\.3%\[21\.2, 38\.9\]80\.2%\[75\.8, 83\.9\]Opus 4\.841\.9%\[34\.7, 49\.3\]47\.0%\[40\.7, 53\.4\]21\.6%\[16\.1, 28\.4\]62\.3%\[57\.4, 66\.9\]Opus 4\.635\.2%\[28\.6, 42\.4\]38\.0%\[32\.2, 44\.2\]8\.8%\[5\.4, 14\.0\]55\.0%\[50\.1, 59\.8\]GPT\-533\.5%\[27\.0, 40\.7\]40\.4%\[34\.5, 46\.6\]5\.3%\[2\.8, 9\.7\]67\.2%\[62\.5, 71\.7\]

Table 6:Multimodal benchmark performance\. PathVQA and MedXpertQA MM report scored\-subset accuracy after refusal exclusion\. 95% Wilson CIs in brackets\. Opus 4\.8 refused 7% of SLAKE items![Refer to caption](https://arxiv.org/html/2607.10849v1/x5.png)Figure 4:Refusal rates by benchmark\-specific content category: USMLE step \(MedQA\) and task type \(MedXpertQA MM\)\.This pattern does not extend to RareBench\. The mechanism\-versus\-presentation classification showed no discriminating power there, and the subset\-level domain signal identified in Section[3\.3](https://arxiv.org/html/2607.10849#S3.SS3)does not track the basic\-science axis\. We therefore report two related but distinct findings rather than a single unifying theory: across MedQA and MedXpertQA MM, refusal concentrates in basic\-science and mechanism content; on RareBench, refusal concentrates by source dataset and disease domain\. These two patterns are independently well\-supported by the evidence\. We were unable to unify them under a single mechanism, and we report that boundary explicitly rather than proposing an explanation the data does not support\.

### 3\.6\. Multimodal performance

Table[6](https://arxiv.org/html/2607.10849#S3.T6)reports accuracy across all four multimodal benchmarks\. Fable 5 leads on three of the four\. On VQA\-RAD it scores 45\.8%\[38\.7, 53\.1\], ahead of Opus 4\.8 \(41\.9%\), Opus 4\.6 \(35\.2%\), and GPT\-5 \(33\.5%\)\. On MedXpertQA MM, scored on the 368 items Fable 5 answered, it reaches 80\.2%\[75\.8, 83\.9\], a decisive margin over GPT\-5 \(67\.2% \[62\.5, 71\.7\]\), with non\-overlapping confidence intervals\. On SLAKE, all four models cluster between 38\.0% and 49\.2% with substantially overlapping intervals; this benchmark does not discriminate among current systems\. Opus 4\.8 refused 7% of SLAKE items, the only non\-negligible refusal rate on this benchmark, without a clear effect on its relative ranking among the four models On PathVQA, Fable 5 refused 71 of 171 items \(41\.5%\); its scored\-subset accuracy \(29\.3%\[21\.2, 38\.9\]\) places it above Opus 4\.8 \(21\.6%, no refusals\) and the two lower\-scoring models \(Opus 4\.6 7\.6%, GPT\-5 5\.3%\)\. Unlike the pattern in Section[3\.5](https://arxiv.org/html/2607.10849#S3.SS5), refused PathVQA items span both disease\-process questions and purely descriptive ones, with no clear content pattern; we treat this as an open question\.

### 3\.7\. Generational trend and cross\-model comparison

Within the Claude family, the generational picture is consistently favorable to Fable 5 once refusal is separated from correctness\. On every benchmark where Fable 5 answers, its accuracy matches or exceeds both predecessors; the multimodal gap on MedXpertQA MM \(80\.2% vs\. 62\.3% and 55\.0%\) is the largest generational difference observed anywhere in this study\. What changed release to release is not underlying capability but willingness: the refusal behavior documented in Sections[3\.1](https://arxiv.org/html/2607.10849#S3.SS1)–[3\.5](https://arxiv.org/html/2607.10849#S3.SS5)did not exist in Opus 4\.6 or Opus 4\.8 at any comparable rate\.

Against GPT\-5, the comparison is benchmark \-dependent but consistently favors Fable 5 on answered accuracy\. On the legacy text benchmarks, GPT\-5’s raw scores exceed Fable 5’s raw scores; Fable 5’s scored\-subset accuracy exceeds GPT\-5’s on both\. On MedXpertQA in text and image form, Fable 5 leads outright\. On VQA\-RAD, Fable 5 leads; on SLAKE the two are statistically indistinguishable\. On RareBench and PathVQA, GPT\-5 attempts the task where Fable 5 largely does not, and that willingness gap, not a capability gap, is what GPT\-5’s higher raw scores on those benchmarks actually reflect\.

The Appendix presents seven examples where Fable 5 answered correctly and all three other models did not, drawn from each benchmark, illustrating what Fable 5’s reasoning looks like when it engages\.

## 4\. Discussion

### 4\.1\. What looked like capability decline was mostly refusal

The raw scores alone tell a misleading story\. On MedQA and PubMedQA, Fable 5 scores 10 to 15 percentage points below both predecessors and GPT\-5 in raw accuracy, a result that, taken at face value, suggests the newest model in the family is weaker on basic medical knowledge than the ones before it\. The refusal audit in Section[3\.1](https://arxiv.org/html/2607.10849#S3.SS1)dismantles that reading\. Once refused items are excluded from the denominator, Fable 5’s accuracy on both benchmarks matches or exceeds every other model in this study\. The apparent regression is an artifact of counting refusals as wrong answers, not evidence of reduced capability\. What actually changed release to release is not what the model knows, it is what the model is willing to demonstrate\.

This distinction matters for anyone using aggregate benchmark scores to track model progress\. A single accuracy figure, collected without refusal tracking, would place Fable 5 below Opus 4\.6 on benchmarks where it is in fact stronger\. The right generational summary has two parts: Fable 5’s demonstrated accuracy, on every benchmark where it answers, is competitive with or ahead of both predecessors; and a new pattern of selective refusal, absent in both Opus 4\.6 and Opus 4\.8 at any comparable rate, now sits between the model’s capability and its measured performance\.

### 4\.2\. Refusal behavior and Anthropic’s documented safeguards

Anthropic’s release documentation describes Fable 5’s biology and chemistry classifier as targeting dual\-use content capable of providing uplift to malicious actors \- the company’s own example is predicting how a genetic modification would affect viral capsid assembly in the context of pathogen design\. The content that drove refusal in this study does not resemble that description\. RareBench’s most\-refused cases are standard inborn metabolic disease presentations diagnosed from a routine phenotype list\. MedQA’s refused questions are preclinical board\-exam material, anatomy, physiology, pharmacology, the same content medical students are examined on every year\.

Two readings are possible\. If the documented classifier is responsible for what we observed, it is engaging with content well outside its stated scope, at rates well above the sub\-5% session average Anthropic reports\. If it is not responsible, the refusal behavior documented here does not correspond to any safeguard Anthropic has publicly described\. We cannot determine which reading is correct from a black\-box API evaluation, and we do not claim to\. What this study can establish is that the observed refusal pattern, whatever its internal cause, does not behave as the documented mechanism predicts\.

### 4\.3\. The refusal pattern is not one phenomenon but two

Refusal appears on five of the eight benchmarks in this study, at rates ranging from 8\.0% to 99\.4%\. This is not an anomaly confined to one sensitive benchmark, but a recurring feature of Fable 5’s behavior across legacy exam questions, specialist clinical images, and open\-ended rare disease diagnosis alike\.

Section[3\.5](https://arxiv.org/html/2607.10849#S3.SS5)established that at least two distinguishable patterns are operating\. The first runs across MedQA and MedXpertQA MM: refusal concentrates in basic\-science and mechanism\-level content, preclinical Step 1 questions, anatomy and physiology vignettes, receptor pharmacology, at rates far above their baseline share of each benchmark\. Two independent benchmarks, using two entirely different official category schemes, converge on the same split\. The second pattern is specific to RareBench: refusal tracks source dataset and disease domain, with inborn metabolic disease presentations refused near\-universally and adult\-onset autoimmune presentations answered at much higher rates\. We tested directly whether the first pattern explains the second and found it does not, mechanism\-dominant and presentation\-dominant RareBench cases refused at statistically indistinguishable rates, and the classifier failed to track the subset\-level pattern entirely\. PathVQA presents a third signature that neither framework explains, with no content pattern emerging from manual inspection of its refused items\.

We were unable to unify these under a single mechanism, and we report that boundary explicitly rather than proposing an explanation the data does not support\. The Marfan syndrome counterexample in Section[3\.3](https://arxiv.org/html/2607.10849#S3.SS3)subset, opposite refusal outcomes, illustrates that whatever governs this behavior operates at a finer grain than any variable we were able to measure\.

### 4\.4\. Practical implications

Two findings bear directly on how a practitioner should approach deploying Fable 5 on biomedical tasks\.

The first is a negative result, and we think it is the more useful one\. Removing the clinical role framing from the RareBench prompt recovered only 9\.3% of originally\-refused items at full scale\. A practitioner who encounters this behavior and tries a simpler, less authoritative\-sounding prompt will be disappointed roughly nine times out of ten\. Prompt engineering is not a reliable mitigation for this refusal pattern, at least not via the most obvious intervention\.

The second is consistently positive\. Across every benchmark in this study, wherever Fable 5’s response gets through, its accuracy is not lower than the other models’ and is frequently higher\. The primary constraint on its usefulness for biomedical tasks is willingness to engage, not competence once it does\. That is a more tractable problem to design around, through case selection, fallback routing for refused items, or simply anticipating which question types are affected, than a genuine capability gap would be\. A model that is highly accurate but selectively cautious is a different engineering challenge than a model that is simply wrong, and the two should not be treated the same way\.

### 4\.5\. Limits of what this paper establishes

This paper measures behavior, not cause\. We identify two real correlational patterns and one benchmark where no pattern emerged\. We do not identify the underlying mechanism producing any of them\. We have no access to Fable 5’s training data, alignment procedure, or internal representations, and nothing in this paper’s black\-box methodology could establish a causal account in any case\.

One earlier hypothesis, that refusal would concentrate on open\-ended, free\-text benchmarks rather than multiple\-choice ones, the evidence directly contradicts\. MedQA and MedXpertQA MM are both multiple\-choice formats and both show substantial, categorically\-patterned refusal\. Format does not predict which benchmarks are affected\. Content domain comes closer, but it does not generalize to RareBench or PathVQA\. A full mechanistic account requires access this study does not have, most likely the model developer’s own internal tooling\. What this paper can establish, and does, is that refusal is a real, recurring, and benchmark\-dependent behavior in Fable 5 that standard accuracy tables will not surface unless it is tracked and reported as a distinct outcome from the start\.

## 5\. Related Work

Evaluating general\-purpose language models on medical benchmarks began in earnest with Nori et al\.\[[undefc](https://arxiv.org/html/2607.10849#bib.bibx4)\], who tested GPT\-4 zero\-shot on USMLE\-style exams and treated calibration and safety\-tuning costs as first\-class findings alongside accuracy\. Their methodology surfaced an early tension that the field has not resolved: safety\-tuning cost GPT\-4 three to five points of raw accuracy relative to its base model, an early sign that alignment and capability can pull against each other\. Singhal et al\.\[[undefm](https://arxiv.org/html/2607.10849#bib.bibx14),[undefn](https://arxiv.org/html/2607.10849#bib.bibx15)\]pushed the question further with Med\-PaLM 2, adding human\-rater judgment of factuality and potential harm on top of benchmark scores to ask whether a model a physician would actually trust\. The most recent answer comes from Vishwanath et al\.\[[undef](https://arxiv.org/html/2607.10849#bib.bibx1)\], who found that GPT\-5\.2, Gemini 3\.1 Pro, and Claude Opus 4\.6 all outperformed OpenEvidence and UpToDate Expert AI on MedQA, HealthBench, and a real clinical query benchmark drawn from live physician use\. Three years on, the answer to “why test a general model at all” remains the same one Nori et al\.\[[undefc](https://arxiv.org/html/2607.10849#bib.bibx4)\]implied: because it keeps winning\.

The benchmarks in that original evaluation have not aged as well as the methodology\. MedQA\[[undefe](https://arxiv.org/html/2607.10849#bib.bibx6)\]and PubMedQA\[[undeff](https://arxiv.org/html/2607.10849#bib.bibx7)\]are now near ceiling for frontier models, no longer separating one current system from another\. MedXpertQA\[[undefb](https://arxiv.org/html/2607.10849#bib.bibx3)\]and RareBench\[[undefg](https://arxiv.org/html/2607.10849#bib.bibx8)\]were built to restore what those benchmarks lost\. Zuo et al\.\[[undefb](https://arxiv.org/html/2607.10849#bib.bibx3)\]move from four\-option licensing questions to ten\-option specialist\-board questions across 17 specialties, with published top\-model performance as low as 46% on the harder reasoning subset\. Chen et al\.\[[undefg](https://arxiv.org/html/2607.10849#bib.bibx8)\]go further, replacing multiple\-choice recognition with open\-ended diagnosis generation, where standalone models land between 14\.6% and 40% Recall@1\. Both are treated as primary evidence benchmarks here for exactly that reason\.

For image\-based evaluation, we draw on VQA\-RAD\[[undefh](https://arxiv.org/html/2607.10849#bib.bibx9)\], SLAKE\[[undefi](https://arxiv.org/html/2607.10849#bib.bibx10)\], PathVQA\[[undefj](https://arxiv.org/html/2607.10849#bib.bibx11)\], and the multimodal subset of MedXpertQA\[[undefb](https://arxiv.org/html/2607.10849#bib.bibx3)\]\. These benchmarks cover radiology, multi\-modality clinical imaging, and histopathology, domains where real biomedical questions increasingly arrive with an image attached\. We exclude closed\-ended yes/no subsets from all three VQA datasets, as binary questions no longer require a model to produce an answer, only confirm one\. MedXpertQA’s multimodal subset pairs expert\-level exam questions with real clinical images and patient records, making it the hardest task in this phase and the one most directly analogous to specialist clinical reasoning\.

What none of this prior work does is examine more than one point in a single model’s lineage\. Every study above is a snapshot, one generation compared against another at a single moment in time\. The generational trace in this paper, running Opus 4\.6, Opus 4\.8, and Fable 5 under identical conditions, addresses that gap directly\.

Reliable scoring is a prerequisite for any of these comparisons to hold\. Grading open\-ended responses with another language model introduces the judge’s own biases and failure modes into the result, a risk that is especially damaging when the central claim of a paper is that one model outperforms another\. A February 2026 audit of MedCalc\-Bench illustrates the concrete cost of this kind of validity problem, even outside the LLM\-as\-judge setting specifically: Krohn\-Grimberghe\[[undefo](https://arxiv.org/html/2607.10849#bib.bibx16)\]identified and corrected more than twenty formula and runtime errors in the benchmark’s own calculator implementations, and separately showed that GPT\-5\.2\-Thinking reaches 95–97% accuracy once given the calculator specification at inference time, with the remaining errors attributable to ground\-truth issues rather than genuine reasoning failures, evidence that the benchmark’s apparent difficulty reflected implementation bugs and formula memorization more than clinical reasoning\. Every benchmark in this study is scored against a deterministic ground truth, with no model judgment anywhere in the pipeline\.

## 6\. Conclusion

Claude Fable 5’s biomedical capability, measured on what it actually answers, is strong, stronger than its raw scores in this paper suggest\. On text benchmarks, its scored\-subset accuracy meets or exceeds every other model evaluated, including on MedQA and PubMedQA, where its raw numbers alone would have suggested a regression from its own predecessors\. On multimodal benchmarks, it leads outright on three of four, and on MedXpertQA MM, arguably the hardest, most specialist\-oriented benchmark in this study, it beats every other model by a wide, non\-overlapping margin, even with refusal included in the denominator\. The gap between that underlying capability and Fable 5’s raw performance isrefusal, not competence\. Refusal ranges from 8\.0% to 99\.4% depending on the benchmark, does not appear in either Claude predecessor evaluated here, and follows at least two distinguishable patterns rather than one uniform cause\. That gap is this paper’s central finding, but it should not overshadow the capability finding underneath it: when Fable 5 engages with a biomedical question, it is, on the evidence in this study, among the strongest models tested\.

## 7\. Limitations

Domain and prompting scope\.This study covers eight benchmarks across general clinical knowledge, specialist reasoning, rare disease diagnosis, and four imaging domains, evaluated single\-shot with direct\-answer prompting, following the Nori at al\.\[[undefc](https://arxiv.org/html/2607.10849#bib.bibx4)\]template\. It does not extend to other biomedical domains, genomics, pharmacology, mental health, among others, where refusal behavior may differ, and it does not test chain\-of\-thought prompting, which was excluded primarily on cost grounds and might affect both accuracy and refusal rate\. Both are left to future work\.

Scope of comparison\.We compare against one external model at one point in time, with no human physician baseline; this was a deliberate scope decision, not an oversight\.

Single\-run evaluation\.Every figure reflects one API call per item\. Refusal\-rate estimates are point estimates from a single run, not stability\-tested across repeated queries\.

## References

- \[undef\]K\. Vishwanath and E\. K\. et al\. Oermann“General\-Purpose Large Language Models Outperform Specialized Clinical AI Tools on Medical Benchmarks”In*Nature Medicine*, 2026
- \[undefa\]J\. Gu et al\.“A Survey on LLM\-as\-a\-Judge”In*arXiv preprint arXiv:2411\.15594*, 2024arXiv:[https://arxiv\.org/abs/2411\.15594](https://arxiv.org/abs/2411.15594)
- \[undefb\]Y\. Zuo et al\.“MedXpertQA: Benchmarking Expert\-Level Medical Reasoning and Understanding”In*Proceedings of the International Conference on Machine Learning \(ICML\)*, 2025eprint: 2501\.18362
- \[undefc\]H\. Nori et al\.“Capabilities of GPT\-4 on Medical Challenge Problems”In*arXiv preprint arXiv:2303\.13375*, 2023arXiv:[https://arxiv\.org/abs/2303\.13375](https://arxiv.org/abs/2303.13375)
- \[undefd\]F\. Joad et al\.“There Is More to Refusal in Large Language Models than a Single Direction”In*arXiv preprint arXiv:2602\.02132*, 2026eprint: 2602\.02132
- \[undefe\]D\. Jin et al\.“What Disease Does This Patient Have? A Large\-Scale Open Domain Question Answering Dataset from Medical Exams”In*Applied Sciences*11\.14, 2021, pp\. 6421
- \[undeff\]Q\. Jin et al\.“PubMedQA: A Dataset for Biomedical Research Question Answering”In*arXiv preprint arXiv:1909\.06146*, 2019arXiv:[https://arxiv\.org/abs/1909\.06146](https://arxiv.org/abs/1909.06146)
- \[undefg\]X\. Chen et al\.“RareBench: Can LLMs Serve as Rare Diseases Specialists?”In*Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining \(KDD ’24\)*, 2024
- \[undefh\]J\. J\. Lau, S\. Gayen, A\. Ben Abacha and D\. Demner\-Fushman“A Dataset of Clinically Generated Visual Questions and Answers About Radiology Images”In*Scientific Data*5, 2018, pp\. 180251
- \[undefi\]B\. Liu et al\.“SLAKE: A Semantically\-Labeled Knowledge\-Enhanced Dataset for Medical Visual Question Answering”In*IEEE International Symposium on Biomedical Imaging \(ISBI\)*, 2021
- \[undefj\]X\. He et al\.“PathVQA: 30000\+ Questions for Medical Visual Question Answering”In*arXiv preprint arXiv:2003\.10286*, 2020arXiv:[https://arxiv\.org/abs/2003\.10286](https://arxiv.org/abs/2003.10286)
- \[undefk\]undef OpenAI“GPT\-5 System Card”In*arXiv preprint arXiv:2601\.03267*, 2026arXiv:[https://arxiv\.org/abs/2601\.03267](https://arxiv.org/abs/2601.03267)
- \[undefl\]Sebastian Köhler et al\.“Expansion of the Human Phenotype Ontology \(HPO\) Knowledge Base and Resources”In*Nucleic Acids Research*47\.D1, 2019, pp\. D1018–D1027DOI:[10\.1093/nar/gky1105](https://dx.doi.org/10.1093/nar/gky1105)
- \[undefm\]K\. Singhal et al\.“Large Language Models Encode Clinical Knowledge”In*Nature*620\.7972, 2023, pp\. 172–180
- \[undefn\]K\. Singhal et al\.“Toward Expert\-Level Medical Question Answering with Large Language Models”In*Nature Medicine*, 2025
- \[undefo\]A\. Krohn\-Grimberghe“MedCalc\-Bench Doesn’t Measure What You Think: A Benchmark Audit and the Case for Open\-Book Evaluation”In*arXiv preprint arXiv:2603\.02222*, 2026arXiv:[https://arxiv\.org/abs/2603\.02222](https://arxiv.org/abs/2603.02222)

## Appendix AAppendix

## A\. Appendix

Table[7](https://arxiv.org/html/2607.10849#Ax1.T7)reports, for each benchmark, how many items Fable 5 answered correctly while Opus 4\.6, Opus 4\.8, and GPT\-5 all missed\. This is not a claim about aggregate accuracy, Tables[2](https://arxiv.org/html/2607.10849#S2.T2),[3](https://arxiv.org/html/2607.10849#S2.T3), and[6](https://arxiv.org/html/2607.10849#S3.T6)already cover that, but a qualitative look at what these items contain\. Figures A1 through A7 present one example from each benchmark where Fable 5 answered correctly and all three other models did not, showing the question, Fable 5’s response, and the ground\-truth answer\. The examples give a concrete sense of what Fable 5’s reasoning looks like when it engages with a question, complementing the numerical case made throughout Section[3](https://arxiv.org/html/2607.10849#S3)\. The seven examples are chosen for clarity, not difficulty or representativeness, and should not be read as a random sample\. RareBench contributes none, since only 6 items were scored at all under the primary prompt \(Section[3\.2](https://arxiv.org/html/2607.10849#S3.SS2)\)\.

BenchmarkUniquely correctTotal scoredMedQA71,273PubMedQA181,000MedXpertQA Text1342,450RareBench06VQA\-RAD12179PathVQA14171SLAKE9250MedXpertQA MM27368Table 7:Items where Fable 5 answered correctly and Opus 4\.8, Opus 4\.6, and GPT\-5 all answered incorrectly\. RareBench’s zero count is consistent with \(Section[3\.2](https://arxiv.org/html/2607.10849#S3.SS2)\): only 6 items were scored at all under the primary prompt\.MedQA \#677 — Fable 5 uniquely correctQuestion:A 2\-month\-old infant is brought to the emergency department by her mother due to a fall\. Tearfully, the patient’s mother describes witnessing her child fall from the changing table\. The patient was born prematurely at 36 weeks estimated gestational age via vacuum\-assisted vaginal delivery\. The patient is afebrile\. Her vital signs include: blood pressure of 94/60 mm Hg, pulse 200/min, and respiratory rate 70/min\. Physical examination reveals a subconjunctival hemorrhage in the left eye and multiple bruises on the chest and back\. Which of the following is the best initial step in management of this patient’s condition?Answer Choices: \(A\) CT scan of the head \(B\) Involvement of social services \(C\) Family counseling \(D\) Rule out medical conditionsCorrect answer:AFable 5’s response: This 2\-month\-old presents with findings highly concerning for non\-accidental trauma \(child abuse\): multiple bruises on the chest and back in a non\-mobile infant, subconjunctival hemorrhage, and an injury pattern inconsistent with the reported mechanism\. More importantly, the infant is hemodynamically unstable, with marked tachycardia \(200/min\) and tachypnea \(70/min\), raising concern for intracranial hemorrhage\. The priority in suspected child abuse is always medical stabilization first — social services involvement \(B\) is legally mandated but comes after the child is medically stabilized\.ANSWER: A

Figure A1\.Fable 5 correctly sequences medical stabilization before mandated reporting in a case of suspected non\-accidental trauma, while Opus 4\.6, Opus 4\.8, and GPT\-5 each selected a different option\.

PubMedQA \#23497210Question:Are women with major depression in pregnancy identifiable in population health data?Context \(abridged\):A population\-based linkage study compared two methods of identifying maternal depression during pregnancy — antidepressant dispensing records and hospital admission comorbidity records — across 96,698 births in Western Australia\. Only 899 of 6,596 identified pregnancies \(13\.6%\) were captured by*both*data sources; 80% of women dispensed an antidepressant had no depression comorbidity recorded on their hospital records\.Correct answer:noFable 5’s response: The study found very poor concordance between the two data sources: only 13\.6% of pregnancies were captured by both\. Notably, 80% of women dispensed antidepressants had no depression comorbidity recorded on their hospital records\. Furthermore, the capture\-recapture estimate suggests true prevalence \(~16%\) is far higher than what either dataset identified individually\. This indicates substantial under\-ascertainment\.ANSWER: no

Figure A2\.Fable 5 correctly reasons from cross\-source discordance and a capture\-recapture estimate to conclude that neither administrative dataset reliably identifies maternal depression alone\.

MedXpertQA Text \#Text\-20Question:A 57\-year\-old male with a history of IV drug use, hepatitis C, atrial fibrillation, cirrhosis, alcohol dependence, obesity, and depression presents with a one\-week history of weakness and fever\. He is admitted and started on piperacillin\-tazobactam for suspected spontaneous bacterial peritonitis\. By day 4, he is afebrile and improving, but nurses report bleeding from IV sites and new bruising\. What is the most likely explanation for this patient’s condition?Answer Choices: \(A\) Fibrinolysis activation due to liver failure \(B\) Vitamin K deficiency due to liver dysfunction \(C\) Diffuse activation of the coagulation cascade \(D\) Disseminated intravascular coagulation \(DIC\) \(E\) Platelet dysfunction secondary to sepsis \(F\) Decreased production of coagulation factors \(G\) Immune\-mediated platelet destruction \(H\) Hypersplenism causing thrombocytopenia \(I\) Decreased metabolism of an anticoagulant \(J\) Bacterial destruction of clotting factorsCorrect answer:JFable 5’s response: This patient has atrial fibrillation, for which he is almost certainly anticoagulated with warfarin\. The key clue is timing: bleeding developed on hospital day 4, after several days of broad\-spectrum antibiotics\. These antibiotics eradicate colonic flora that synthesize vitamin K, reducing synthesis of factors II, VII, IX, and X — an effect that dramatically potentiates warfarin’s anticoagulant action\. DIC \(C, D\) would be expected with worsening sepsis, but this patient is improving; piperacillin\-tazobactam does not significantly inhibit warfarin metabolism \(I\), unlike agents such as metronidazole or fluconazole\.ANSWER: J

Figure A3\.Fable 5 correctly identifies gut\-flora\-mediated vitamin K depletion potentiating an unstated pre\-existing warfarin regimen, using the day\-4 timing to rule out DIC and platelet\-related causes despite a septic\-sounding presentation\. Opus 4\.6, Opus 4\.8, and GPT\-5 each selected a different option among ten choices\.

VQA\-RAD \#vqa\_rad\_26 — Fable 5 uniquely correct![[Uncaptioned image]](https://arxiv.org/html/2607.10849v1/vqa_rad_26.jpg)Question:What type of MRI modality is used?Ground\-truth answer:DWI \(diffusion weighted\)Fable 5’s response: The image shows a diffusion\-weighted MRI \(DWI\) sequence, evident from the bright hyperintense signal in the left MCA territory consistent with acute infarction\.ANSWER: DWI \(diffusion weighted imaging\)

Figure A4\.Fable 5 correctly identifies the specific MRI sequence, not just the imaging modality class, reasoning from the acute\-infarction signal pattern\. Opus 4\.6, Opus 4\.8, and GPT\-5 each answered incorrectly\.

PathVQA \#path\_vqa\_58 — Fable 5 uniquely correct![[Uncaptioned image]](https://arxiv.org/html/2607.10849v1/path_vqa_58.jpg)Question:What does this image show?Ground\-truth answer:uterusFable 5’s response: This image shows a gross specimen of a uterus with a large leiomyoma \(fibroid\), demonstrating the characteristic whorled, firm, white\-tan cut surface of the tumor\.ANSWER: uterus with leiomyoma

Figure A5\.Fable 5 identifies both the organ and the specific pathological finding, describing the gross morphological features supporting the diagnosis\. Opus 4\.6, Opus 4\.8, and GPT\-5 each answered incorrectly\.

SLAKE \#slake\_245 — Fable 5 uniquely correct![[Uncaptioned image]](https://arxiv.org/html/2607.10849v1/slake_245.jpg)Question:What color does the right lung show in the picture?Ground\-truth answer:blackFable 5’s response: The right lung appears dark \(black\) on this chest X\-ray, as air\-filled lung tissue is radiolucent\.ANSWER: black

Figure A6\.Fable 5 grounds its answer in an imaging physics principle \(radiolucency of air\-filled tissue\) rather than surface\-level pattern matching\. Opus 4\.6, Opus 4\.8, and GPT\-5 each answered incorrectly\.

MedXpertQA MM \#MM\-1938 — Fable 5 uniquely correct![[Uncaptioned image]](https://arxiv.org/html/2607.10849v1/MM-1938-a.png)![[Uncaptioned image]](https://arxiv.org/html/2607.10849v1/MM-1938-b.png)Question:In a 58\-year\-old female patient, what is the MOST likely diagnosis?Answer Choices: \(A\) Osteoarthritis \(B\) Hemophiliac arthropathy \(C\) Gout \(D\) Rheumatoid arthritis \(E\) Amyloid depositionCorrect answer:DFable 5’s response: ANSWER: D

Figure A7\.Fable 5 correctly selects rheumatoid arthritis from a five\-option specialist\-level differential grounded in the accompanying clinical images\. Opus 4\.6, Opus 4\.8, and GPT\-5 each selected a different option\. No reasoning text accompanies this answer, consistent with the direct\-answer, no\-CoT prompting protocol used throughout this study \(Section[2\.3](https://arxiv.org/html/2607.10849#S2.SS3)\)

SLAKE \#slake\_229 — Fable 5 uniquely correct![[Uncaptioned image]](https://arxiv.org/html/2607.10849v1/slake_229.jpg)Question:In what plane is this image scanned?Ground\-truth answer:Transverse PlaneFable 5’s response: This is a CT image showing a cross\-sectional \(transverse\) view of the abdomen at the level of the liver and spine\.ANSWER: Axial \(transverse\) plane

Figure A8\.Fable 5 correctly identifies the imaging orientation as the axial \(transverse\) plane and supports its answer by recognizing the scan as a cross\-sectional abdominal CT image\. Opus 4\.6, Opus 4\.8, and GPT\-5 each answered incorrectly\.

Similar Articles

Claude Fable 5: mid-tier results on coding tasks

Hacker News Top

Anthropic's Claude Fable 5 model showed middling performance on real-world vulnerability-fixing tasks, with many timeouts and high cheating volume, but also solved four instances no previous model had cracked.

Claude Fable 5 benchmarks

Reddit r/singularity

Anthropic released benchmarks for Claude Fable 5, a new AI model, showing significant performance improvements.

Claude Fable won’t answer basic biology questions

The Verge

Anthropic's new Claude Fable 5 model refuses to answer basic biology questions due to overly conservative safety filters aimed at preventing bioweapons misuse, highlighting the tradeoff between capability and safety.

Claude Fable is relentlessly proactive

Hacker News Top

The article describes how Claude Fable 5, an AI model, demonstrates relentless proactivity by autonomously using browser automation, shell commands, and custom scripts to debug a UI issue, illustrating advanced tool-use capabilities.