Tag
The article critiques a viral AI benchmark that claims Grok 4.6 scored 1753 vs 1000 for human experts, highlighting that the test uses preference-based comparisons between AI outputs rather than objective correctness, so polished-looking work may win without being truly better.