@elonmusk: Grok 4.5 reaches #1 position on Long-Horizon Terminal-Bench

X AI KOLs Timeline Models

Summary

Elon Musk announces that Grok 4.5 has achieved the #1 position on the Long-Horizon Terminal-Bench benchmark, surpassing previous models.

Grok 4.5 reaches #1 position on Long-Horizon Terminal-Bench
Original Article
View Cached Full Text

Cached at: 07/14/26, 02:26 PM

Grok 4.5 reaches #1 position on Long-Horizon Terminal-Bench

tetsuo (@tetsuoai): The Long-Horizon Terminal-Bench paper landed around May and concluded that the results showed headroom for improvement. The best of the 15 models they tested finished seven of the 46 tasks, and the mean across all models was about two. That ceiling is what fifth place looks like

Similar Articles

@elonmusk: Grok 4.7 moves up in ranking

X AI KOLs Following

Grok 4.7 has moved up in the SWE-Together leaderboard rankings after its weak spots were identified and fixed, leading to an audit and update of all model trials.

@elonmusk: Grok Build

X AI KOLs Timeline

Grok 4.5 with Grok Build achieved #1 on the SWE-Atlas-QnA benchmark with a score of 84, matching GPT-5.6 Codex and outperforming other coding setups.