@_xjdr: to get a better sense of the gap between OSS - frontier i find it helpful to think of dsv4-flash as a sonnet tier model…
Summary
Discussion of open-source model tiers, comparing DSV4-flash to Sonnet 5 and GLM 5.2 to Opus 4.8, with a prediction of a fable-tier model by end of year.
Similar Articles
@andrewchen: honestly pretty incredible DSV4 Flash 0731 versus Opus 4.6:
Andrew Chen shares excitement about testing DeepSeek V4 Flash 0731 on dual NVIDIA DGX Sparks, comparing it to Opus 4.6.
With release of Deepseek V4 I wanted see how the model sizes are trending over time. Open source models are constantly getting smaller and better. The trend is that by this time next year, we probably will have Opus 4.5 level models on consumer grade laptops (sounds unlikely?!).
The author analyzes model size and performance trends following Deepseek V4 Flash, suggesting that open-source models are shrinking in size while improving, and predicts Opus 4.5-level models could run on consumer laptops within a year.
Claude Sonnet 5 is out and the gap with Opus 4.8 is smaller than I expected
Anthropic released Claude Sonnet 5, which achieves benchmark scores very close to Opus 4.8 at a significantly lower price, making it a compelling option for agentic tasks despite potential real-world gaps.
@FuckAnthropic: Conducted a comparative analysis. Overall, DeepSeek V4 Flash-0731 is roughly a model at the level between Opus 4.7 and 4.8, entering the frontier Agent model competition with a minimal activation scale, and at about 1/12 to 1/60 of the token cost to enter the frontier Ag…
The author's comparative analysis concludes that DeepSeek V4 Flash-0731 achieves Opus 4.7–4.8 level performance with an extremely small activation scale, entering the frontier agent model tier at a very low token cost. It surpasses GLM-5.2 overall, but its shortfalls remain difficult repository-level coding and long-horizon engineering.
@danshipper: This is wrong, it’s the same model But it does fall back to Opud 4.8 slightly more, so the benchmarks are measuring a m…
Dan Shipper argues that the Fable 5 model is not nerfed but falls back to Opus 4.8 more often, causing mixed benchmark results, contrary to claims of severe degradation.