@elonmusk: Grok 4.7 moves up in ranking
Summary
Grok 4.7 has moved up in the SWE-Together leaderboard rankings after its weak spots were identified and fixed, leading to an audit and update of all model trials.
View Cached Full Text
Cached at: 09/24/26, 12:12 AM
Grok 4.7 moves up in ranking
Zhuokai Zhao (@zhuokaiz): After we fixed the weak spots exposed by Grok 4.7 (thank you, Grok), we audited every model we have run on the SWE-Together leaderboard for the same behavior, re-ran every trial that got through, and updated the rows.
Here is what changed.
We scanned the tool calls of all
Similar Articles
@elonmusk: Grok 4.5 reaches #1 position on Long-Horizon Terminal-Bench
Elon Musk announces that Grok 4.5 has achieved the #1 position on the Long-Horizon Terminal-Bench benchmark, surpassing previous models.
@elonmusk: Grok 4.6 ranks #1 on CursorBench for real-world coding
Elon Musk announces Grok 4.6 ranked #1 on CursorBench for real-world coding, outperforming other AI models and showcasing high efficiency.
@elonmusk: Grok model improvement
The updated Grok model (0.5T) is less lazy, more autonomous, and more accurate; improvements are ongoing.
@elonmusk: Grok Build
Grok 4.5 with Grok Build achieved #1 on the SWE-Atlas-QnA benchmark with a score of 84, matching GPT-5.6 Codex and outperforming other coding setups.
@elonmusk: And Grok 4.6 is a significant improvement
Elon Musk announces that Grok 4.6 is a significant improvement over previous versions.