@_jasonwei: Muse Spark 1.1 outperforms GPT-5.6 Sol and Gemini 3.1 on Radiology's Last Exam. We don't beat Fable (yet). And Humans a…
Summary
Muse Spark 1.1 outperforms GPT-5.6 Sol and Gemini 3.1 on Radiology's Last Exam 2.0, a new visual reasoning benchmark for autonomous AI diagnosis in healthcare, though it still lags behind Fable and human radiologists.
View Cached Full Text
Cached at: 07/13/26, 08:05 PM
Muse Spark 1.1 outperforms GPT-5.6 Sol and Gemini 3.1 on Radiology’s Last Exam. We don’t beat Fable (yet). And Humans are still a lot better, but we are working on closing the gap!
Dr. Datta M.D. (Radiology) ✈️ Switzerland @AI4Good (@DrDatta_AIIMS): 🔥Today, we are releasing one of the first visual reasoning benchmarks for autonomous AI diagnosis in healthcare!
🚀Introducing Radiology’s Last Exam 2.0 (RadLE 2.0) from @CRASHLabAI, an uncertainty-aware benchmark for autonomous diagnosis in radiology!
✅In the last few days,
Similar Articles
Meta’s muse spark 1.3 surpassed fable 5 and GPT 5.6 sol
Meta's Muse Spark 1.3 AI model has reportedly surpassed Fable 5 and GPT 5.6 in performance metrics.
@FinanceYF5: Alexandr Wang stated that Muse Spark 1.1 is an industry-competitive agentic and coding model. In multiple agentic benchmarks, it can compete with GPT-5.5 and Opus 4.8. It is now available via the new Meta…
Alexandr Wang stated that Muse Spark 1.1 is an industry-competitive agentic and coding model, capable of competing with GPT-5.5 and Opus 4.8 in multiple benchmarks. It is now available via the Meta Model API and Meta AI.
@_jasonwei: In addition to agents and coding, Muse Spark 1.1 is also really strong at answering health questions, a steadily growin…
Muse Spark 1.1, a new agentic and coding model from Meta, achieves +5% improvement on HealthBench-Pro, outperforming all competitors except Fable and Mythos.
@rohanpaul_ai: Very interesting experiments by @thehypedotnews gpt-6 sol vs grok 4.7 vs gpt-6 astra vs muse spark 1.3 Amazing how litt…
Experiments comparing AI models like GPT-6 Sol, Grok 4.7, GPT-6 Astra, and Muse Spark 1.3 in generating browser artifacts from prompts highlight GPT-6 Sol's efficiency with minimal inference to produce working HTML files.
SemiAnalysis: Gemini 3.8 Flash and Muse Spark 1.3 are two of the most clearly benchmaxxed models we've seen yet.
SemiAnalysis reports that Gemini 3.8 Flash and Muse Spark 1.3 are among the most clearly benchmarked AI models, showcasing strong performance.