@_jasonwei: Muse Spark 1.1 outperforms GPT-5.6 Sol and Gemini 3.1 on Radiology's Last Exam. We don't beat Fable (yet). And Humans a…

X AI KOLs Following Models

Summary

Muse Spark 1.1 outperforms GPT-5.6 Sol and Gemini 3.1 on Radiology's Last Exam 2.0, a new visual reasoning benchmark for autonomous AI diagnosis in healthcare, though it still lags behind Fable and human radiologists.

Muse Spark 1.1 outperforms GPT-5.6 Sol and Gemini 3.1 on Radiology's Last Exam. We don't beat Fable (yet). And Humans are still a lot better, but we are working on closing the gap!
Original Article
View Cached Full Text

Cached at: 07/13/26, 08:05 PM

Muse Spark 1.1 outperforms GPT-5.6 Sol and Gemini 3.1 on Radiology’s Last Exam. We don’t beat Fable (yet). And Humans are still a lot better, but we are working on closing the gap!

Dr. Datta M.D. (Radiology) ✈️ Switzerland @AI4Good (@DrDatta_AIIMS): 🔥Today, we are releasing one of the first visual reasoning benchmarks for autonomous AI diagnosis in healthcare!

🚀Introducing Radiology’s Last Exam 2.0 (RadLE 2.0) from @CRASHLabAI, an uncertainty-aware benchmark for autonomous diagnosis in radiology!

✅In the last few days,

Similar Articles