@rasbt: Cool new open-weight model by Cohere: a new lightweight 30B open-weight model for agentic coding tasks. This one builds…
Summary
Cohere released a new lightweight 30B open-weight model for agentic coding tasks, built on Command A+ with parallel transformer design, showing strong performance on agentic benchmarks like Terminal-Bench and SWE-Bench.
Similar Articles
Deepseek V4 Flash is now ~#2 open weight model to Kimi K3 and >50x cheaper
DeepSeek's new V4 Flash model is reportedly the #2 open-weight model behind Kimi K3, offering strong performance at over 50x lower cost ($0.09/$0.18 per 1M tokens) with solid coding and reasoning capabilities.
Orca-Bench: How Ready Are Language Model Agents for Oncall?
Introduces ORCA-bench, a production-fidelity benchmark for evaluating LLM agents on oncall root cause analysis, finding that even frontier agents achieve only 25.3% accuracy on medium-difficulty tasks.
DeepSeek V4 Flash GA ranks the same as Sonnet 5 and Grok 4.5 on DeepSWE
DeepSeek's V4 Flash GA model matches the performance of Sonnet 5 and Grok 4.5 on the DeepSWE benchmark.
DeepSeek-V4-Flash-0731 unsloth gguf on A100
DeepSeek-V4-Flash-0731 is shown running as an unsloth GGUF quant on a single 40GB A100, with 17.7 tok/s and 6 experts loaded into VRAM, enabling a full agentic coding loop.
I had Claude and OpenAI Codex each write a chess engine from one prompt, then made them play 10 games. 10-0, all checkmates, and Codex lost the identical 24-move game five times
A user had Claude and OpenAI Codex each write a chess engine from a single prompt, then made them play 10 games. Claude won all 10 by checkmate, and Codex repeatedly lost the identical 24-move game.