Tag
The article critiques a viral AI benchmark that claims Grok 4.6 scored 1753 vs 1000 for human experts, highlighting that the test uses preference-based comparisons between AI outputs rather than objective correctness, so polished-looking work may win without being truly better.
Elon Musk announces that Grok 4.6 is optimized for the Grok Build harness, and performance is significantly worse without it.
Aravind Srinivas congratulates SpaceXAI on Grok 4.6, noting that it performs well on the Wide-And-Deep-Research benchmark using the Perplexity Computer harness, and is now available to Pro and Max users.
Grok 4.6 is officially released, with the same pricing as 4.5, enhanced agent and programming capabilities, and it runs longer on complex tasks. The AA Intelligence Index ties GPT-5.6 Sol, with input at $2 per million tokens and output at $6 per million tokens.
A hands-on field guide to Grok 4.6, highlighting its speed, dense communication style, and effective prompting patterns for coding and knowledge work.
Ray Fernando hosts a live broadcast design challenge using Grok 4.6 in Cursor, showcasing AI-assisted design workflows.
xAI releases Grok 4.6, a frontier model focused on long-running agents and ambitious interactive/visual work, matching GPT-5.6 Sol on the Artificial Analysis Intelligence Index and available in Cursor and Grok Build.