@QuixiAI: ChatGPT failed Claude failed Deepseek failed Qwen3.8 passed GLM 5.3 passed Grok passed
Summary
The tweet reports that ChatGPT, Claude, and Deepseek failed a test or benchmark, while Qwen3.8, GLM 5.3, and Grok passed, based on a linked source.
View Cached Full Text
Cached at: 08/29/26, 08:01 AM
ChatGPT failed Claude failed Deepseek failed Qwen3.8 passed GLM 5.3 passed Grok passed https://t.co/q7CbBbPymk
Similar Articles
ChatGPT, Gemini, Claude, Grok Fail Accuracy Test on Election Topics: Forum AI
A study by Forum AI found that major chatbots like ChatGPT, Gemini, Claude, and Grok fail to provide accurate and unbiased election information, with 90% of responses containing errors or bias.
Weekly AI roundup (May 23–30, 2026): Claude Opus 4.8 Fast Mode 3x cheaper, Qwen 3.7 Max beats Claude at half the price, ChatGPT moves into Excel
A comprehensive roundup of major AI releases from May 23–30, 2026, covering price cuts for Claude Opus 4.8 Fast Mode, the launch of Qwen 3.7 Max with competitive pricing, ChatGPT integration into Excel, Gemini 3.5 Flash, Grok Build 0.1, Mistral's Vibe agent, and Hugging Face's robot app store, with analysis on falling inference costs and the battleground shifting to distribution.
@charles_irl: new benchmark just dropped
Andon Labs released a new benchmark testing whether AI models refuse to play a Nazi marching song, finding that Claude Opus 4.8 and GPT 5.5 always refused, Gemini 3.5 Flash refused half the time, and Grok 4.3 almost always played it.
Local Qwen 3.8 27B vs GPT‑5.6 Terra vs Grok 4.6
The article compares three AI models—Qwen 3.8 27B, GPT-5.6 Terra, and Grok 4.6—on their ability to build a Three.js fragrance launch site, detailing their implementation strengths and potential issues.
I tested every frontier model from every AI lab - Claude Fable 5, GPT 5.6 Sol, Kimi K3, GLM 5.3, Qwen 3.8 Max, DS v4 Pro, Grok 4.6 and just 1 made it through.
The author tested multiple frontier AI models on extracting data from large log files, finding that only Claude Fable 5 succeeded by streaming data instead of loading files into memory, highlighting its superior practical intelligence compared to others.