DeepSeek-V4-Flash-0731: When Low is higher than High
Summary
A developer benchmarks DeepSeek-V4-Flash-0731 across four reasoning effort modes (none, low, high, max), finding that Low mode is surprisingly verbose and that OpenRouter has a bug affecting reasoning effort modes.
Similar Articles
DeepSeek V4 Pro 0813 (on OpenRouter)
Simon Willison covers the release of DeepSeek V4 Pro 0813, available via API on OpenRouter, noting the likely open weights and noticeable differences between reasoning levels.
Why no "high" reasoning effort in Qwen 3.8 27b ?
The article questions why the Qwen 3.8 27b model lacks a 'high' reasoning effort mode, highlighting a large gap between the 'medium' and 'xhigh' settings.
The difference between "medium" and "xhigh" reasoning effort for Qwen3.8-27B is actually insane.
The user observes that setting reasoning effort to 'xhigh' in Qwen3.8-27B generates significantly more thinking tokens compared to 'medium', with dramatic differences in token usage.
DeepSeek V4 Flash 0731
DeepSeek V4 Flash 0731 presents its results on the ARC-AGI benchmark, highlighting progress in abstract reasoning for AI models.
Is anyone else finding DeepSeek-V4-Flash unreliable for non-coding tasks?
A user reports that DeepSeek-V4-Flash-0731 is unreliable for non-coding office tasks like summarization and meeting notes, failing at concept extraction and speaker understanding despite strong benchmark scores, while Gemma-4-31B performs better.