@LeonEnglaender: We're just 8 people on our core code team and our 30B-A3B model lands on par with Claude Haiku 4.5 and ahead of NVIDIA'…
Summary
A team of 8 released a 30B-A3B coding model under Apache 2.0 that matches Claude Haiku 4.5 performance and beats NVIDIA's 120B-A12B Nemotron 3 Super on the Artificial Analysis Coding Index.
View Cached Full Text
Cached at: 06/10/26, 05:48 AM
We’re just 8 people on our core code team and our 30B-A3B model lands on par with Claude Haiku 4.5 and ahead of NVIDIA’s 120B-A12B Nemotron 3 Super on the Artificial Analysis Coding Index. Released under Apache 2.0. Very proud of our work & lots more to come! https://t.co/F2wu66R16h
Cohere (@cohere): Introducing Cohere’s first open-source coding model: North Mini Code
Small & efficient, designed for agentic performance and built for community input.
Similar Articles
Open-source models are closing the coding gap with GPT/Claude/Gemini ~1.5x faster than the frontier is advancing, and on decontaminated benchmarks a 27B model already beats Claude Opus 4.8 [live dashboard + analysis]
A live dashboard and statistical analysis shows open-source coding models are closing the gap with closed models at 1.5x the rate, with a 27B model already surpassing Claude Opus on decontaminated benchmarks. Tool-call reliability remains the main bottleneck.
NVIDIA just announced the release of Nemotron 3 Ultra (2 minute read)
Anthropic released Claude Opus 4.5, its most intelligent model, scoring 70 on the Artificial Analysis Intelligence Index and ranking second only to Gemini 3 Pro. It achieves significant gains in coding and agentic tasks while reducing per-token pricing and maintaining strong safety performance.
@xdotli: my friend @xeophon thinks coding is solved here's validation that a 3b model is trained with focus on algo efficiency a…
Nanbeige 4.1, a 3B model, outperforms Qwen3-30b-A3b and Qwen 3.5 4b in coding tasks with focus on algorithmic efficiency, achieving long horizon tasks with 600+ tool calls.
@elliotarledge: Claude Fable 5 [max] on KernelBench-Hard. The main kernel that impressed me was a B200 fp8 GEMM: it HAND WROTE raw SM10…
Claude Fable 5 achieves top results on KernelBench-Hard by hand-writing PTX code for B200 fp8 GEMM, outperforming other models and reaching 44-59% of peak performance on compute-bound shapes.
@jinyuhou0: On popular benchmarks, our 30B model matches systems 20-30x its size (gpt-5.4-xhigh, DeepSeek-V3.2, Kimi-K2.5), while u…
A new 30B model matches systems 20-30x its size on popular benchmarks while using up to 95% fewer reasoning tokens than comparable agentic LLMs, achieved through a learned configurator that decides when and how to reason. Model and code are openly available.