@kimmonismus: Holy, China strikes again: Qwen3.8-Max reportedly worked autonomously for 16 days while costing 80% less than GPT-5.6 S…
Summary
Alibaba announces Qwen3.8-Max, a 2.4T-parameter MoE frontier model with open weights coming next week, claiming autonomous operation for 16 days and significantly lower cost than GPT-5.6 Sol and Claude Fable 5.
View Cached Full Text
Cached at: 08/03/26, 09:38 AM
Holy, China strikes again: Qwen3.8-Max reportedly worked autonomously for 16 days while costing 80% less than GPT-5.6 Sol and 88% less than Claude Fable 5 on output. And its open weight!
Alibaba’s 2.4T-parameter MoE costs $2/M input tokens and $6/M output tokens.
GPT-5.6 Sol: 5/30. Claude Fable 5: 10/50.
Qwen says the model operated autonomously for 16 days, producing 265 commits, 127 PRs and 151 issues through an issue -> code -> test -> repair -> merge loop.
It does not lead every benchmark. But it reaches the frontier range across coding, professional work and computer use, while PaperBench puts it ahead of both Fable 5 and GPT-5.6 Sol at a very good pricing.
Open weights arrive next week.
A model that can economically work for ten days may be more useful than a better model you can afford to run for ten minutes.
Ngl another insane china release.
Qwen (@Alibaba_Qwen): 📢Meet Qwen3.8-Max — our most capable model to date.
Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!🎉
Qwen3.8-Max, a new bar for coding and cowork at 2.4T parameters:
- Autonomous coding: 10+ days of
Similar Articles
Alibaba's Qwen3.7-Max Ran Autonomously for 35 Hours on Unfamiliar Hardware. It Still Kept Getting Better.
Alibaba's Qwen3.7-Max model autonomously optimized a production kernel on unfamiliar T-Head PPU hardware over 35 hours, making 1,158 tool calls and achieving a 10x speedup, demonstrating sustained autonomous agentic behavior without human guidance.
@cline: Qwen3.8-Max is Alibaba’s largest model yet at 2.4T params, and shows a 2% higher benchmark result on Terminal-Bench tha…
Alibaba unveils Qwen3.8-Max, its largest model at 2.4T parameters, showing a 2% higher Terminal-Bench result than Fable 5, with open weights to be released next week.
@rohanpaul_ai: Alibaba released Qwen3.8-Max, a 2.4 trillion parameter model that activates only about 95 bn parameters per token. A th…
Alibaba released Qwen3.8-Max, a 2.4 trillion-parameter sparse MoE model with 95B active parameters per token, 1M token context, and strong agentic and benchmark results, including autonomously coding for days, circuit design, and outperforming rivals on Terminal Bench and PaperBench.
@WEB3_furture: COOL! Someone took the newly released Qwen 3.7-Max, Claude Opus 4.7, and GPT-5.5 for an Agent loop comparison: letting the model write its own Tetris bot, test it, and directly PK after 10 consecutive iterations. Results: Qwen 3.7-Max: +$…
Someone conducted an Agent loop comparison test on Qwen 3.7-Max, Claude Opus 4.7, and GPT-5.5, letting the models write their own Tetris bots and iterate 10 rounds before competing. The results show that Qwen 3.7-Max leads in both performance and cost.
GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B and most today’s low-tier models
GPT-5, which was the best model just a year ago, is now outperformed by Qwen3.6 27B and other current low-tier models, highlighting the rapid pace of AI advancement.