Tag
Cognition's SWE-2 is a new coding model post-trained from Kimi K3 with reinforcement learning, achieving 50.0% on FrontierCode 1.1 Main and offering cost savings of 64% compared to Fable 5.1 while being competitive with GPT-6 Astra. It is now available in Devin Desktop and CLI.
Cognition launches SWE-2, an advanced coding model that achieves competitive performance with Fable 5.1 and GPT-Astra at a fraction of the cost, leveraging novel reinforcement learning to optimize the cost-performance frontier.
Z.AI releases GLM-5.3 Fast, an advanced open-weight AI model optimized for agentic coding and cybersecurity, featuring a 744B-A40B MoE architecture with substantial benchmark improvements.
The article details an open-source reproduction of training a coding model to generate watercolour art using TRL and OpenEnv, with full pipeline artifacts published on Hugging Face for reproducibility.
A mysterious AI model named Ox Alpha has been released on OpenRouter, leading to widespread speculation about its developers, with theories pointing to Chinese AI companies like Z.ai or other entities.
Microsoft introduces MAI-Code-1.1-Flash, a faster, cheaper, and more token-efficient coding model now in production in GitHub Copilot, with improvements in CLI and .NET tasks.
Microsoft's MAI-Code-1.1-Flash coding model is rolling out in GitHub Copilot, adding native vision support and improved coding performance at a 73% lower list price than its predecessor.
A developer enthusiastically recommends KAT Coder 2.5 dev, claiming it is faster, more accurate, and uses fewer tokens than Qwen 3.6 35b a3b, and outperforms Gemma 4 models on their setup, with a GitHub repo containing detailed benchmarks.
Kimi Code releases Kimi K3-256k, a 256k-context version of its flagship K3 coding model, offering reduced quota consumption while maintaining similar performance for most tasks.
Lithos announces its inference engine serving Kimi K2.7 Code, achieving over 1,000 tokens/sec per user on a single 8×B200 node at native precision, 3.4–5.7× faster than major providers.
Meta launched Muse Spark 1.1, a multimodal AI model for agentic coding, competing with OpenAI and Anthropic at a competitive price.
Databricks tested GLM-5.2, an open-source coding model, and found it competes with top closed models like Claude Opus 4.8 on real enterprise code tasks while being cheaper ($1.28/task vs $1.94/task). The evaluation also highlighted Pi, a harness that reduces costs by sending less context per turn.
A developer created a tool that chains a small local model with a larger coding model, automatically offloading VRAM between them to optimize memory usage.
Introduced MTP speculative decoding to the Ornith 35B coding model in FP8 precision, achieving approximately 18% faster inference with minimal extra VRAM.
Cohere Labs releases North Mini Code, a 30B parameter (3B active) open-source coding model under Apache 2.0, optimized for code generation and agentic tasks, capable of running locally with 20GB RAM via 4-bit quantization.
Kimi released the K2.7 Code model and its high-speed version, and announced API pricing. Compared to rival Mimo, it is more expensive and slower.
Open-source model builders have released Ornith-1.0-397B, a top-tier coding model, while larger models like Fable 5 and GPT-5.6 face restrictions.
Kog has open-sourced the Laneformer 2B model, a 2.3B parameter instruction-tuned coding model designed for high-speed decoding, achieving over 3,000 tokens per second by prioritizing latency from the architecture stage.
Jackrong releases Qwopus3.6-27B-Coder-Compat-MTP-GGUF, a GGUF quantization of the Qwopus3.6-27B-Coder model with an expanded chat template for better interoperability with tool-using runtimes and OpenAI-compatible agent frameworks.
GLM 5.2 is a frontier open-source coding model that performs near Claude Opus quality on coding tasks, with excellent tool calling, planning, and local deployment capabilities, at no cost.