Tag
This paper evaluates Claude Fable 5 on eight biomedical benchmarks, finding that despite high refusal rates (8-99.4%), the model achieves superior accuracy when it does answer, highlighting willingness to engage as the primary constraint.