@elonmusk: Grok 4.6 is #1 on healthcare questions
Summary
Grok 4.6 has achieved the top position on MedAgentBench, outperforming GPT-5.6 Sol and Grok 4.5 in real-world healthcare AI tasks.
View Cached Full Text
Cached at: 08/19/26, 06:45 PM
Grok 4.6 is #1 on healthcare questions
X Freeze (@XFreeze): Grok 4.6 just took the #1 spot on MedAgentBench….one of the most interesting benchmarks for real-world agentic healthcare tasks
and Grok now has two generations sitting in the top 3
• #1 Grok 4.6 — ~95.9% • #2 GPT-5.6 Sol — ~94.7% • #3 Grok 4.5 — ~93.4%
MedAgentBench goes
Similar Articles
@elonmusk: Grok 4.6 ranks #1 on CursorBench for real-world coding
Elon Musk announces Grok 4.6 ranked #1 on CursorBench for real-world coding, outperforming other AI models and showcasing high efficiency.
@elonmusk: Grok
Elon Musk highlights that xAI's Grok 4.6 tops RareBench for rare-disease diagnosis, beating Claude Opus 5 at roughly one-third the cost, while DeepSeek's new model underperforms.
@elonmusk: Grok Build
Grok 4.5 with Grok Build achieved #1 on the SWE-Atlas-QnA benchmark with a score of 84, matching GPT-5.6 Codex and outperforming other coding setups.
SpaceXAI's Grok 4.6 Scores 61 on the Artificial Analysis Intelligence Index
SpaceXAI's Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, joining the frontier alongside GPT-5.6 Sol and Claude models, with strong agentic performance at lower cost.
@elonmusk: Grok 4.5 reaches #1 position on Long-Horizon Terminal-Bench
Elon Musk announces that Grok 4.5 has achieved the #1 position on the Long-Horizon Terminal-Bench benchmark, surpassing previous models.