@elonmusk: Grok 4.5 reaches #1 position on Long-Horizon Terminal-Bench
Summary
Elon Musk announces that Grok 4.5 has achieved the #1 position on the Long-Horizon Terminal-Bench benchmark, surpassing previous models.
View Cached Full Text
Cached at: 07/14/26, 02:26 PM
Grok 4.5 reaches #1 position on Long-Horizon Terminal-Bench
tetsuo (@tetsuoai): The Long-Horizon Terminal-Bench paper landed around May and concluded that the results showed headroom for improvement. The best of the 15 models they tested finished seven of the 46 tasks, and the mean across all models was about two. That ceiling is what fifth place looks like
Similar Articles
@elonmusk: Grok 4.6 ranks #1 on CursorBench for real-world coding
Elon Musk announces Grok 4.6 ranked #1 on CursorBench for real-world coding, outperforming other AI models and showcasing high efficiency.
@elonmusk: Grok 4.7 moves up in ranking
Grok 4.7 has moved up in the SWE-Together leaderboard rankings after its weak spots were identified and fixed, leading to an audit and update of all model trials.
@elonmusk: Grok 4.6 reaches #1 on @databricks
Elon Musk announces that Grok 4.6 reached #1 on Databricks, with Ivan Zhou reporting SOTA performance on OfficeQA Pro V2 using Databricks's Genie harness.
@elonmusk: Grok Build
Grok 4.5 with Grok Build achieved #1 on the SWE-Atlas-QnA benchmark with a score of 84, matching GPT-5.6 Codex and outperforming other coding setups.
@elonmusk: Grok Voice is #1!
Elon Musk announces that Grok Voice has reached the number one ranking.