Tag
A developer built a custom benchmark for coding AI agents and found that a 35B-parameter model outperformed a 120B-parameter one when the harness was optimized, highlighting the importance of tailored evaluation over generic specifications.
Qwen3.8-Max-0902 debuts at #1 in Code Arena: WebDev with 1691 points, surpassing Claude Opus 5 by 3 points, and claims the highest-scoring position on the Pareto frontier at $5/MToken.
Proximal has released FrontierSWE v2, an updated ultra-long horizon coding benchmark with expanded tasks and improved methodology, highlighting large performance gaps where Claude Fable 5.1 leads.
GLM-5.3 achieves 310 tokens per second on Databricks inference, leading in both speed and latency, and is the strongest open-source model for coding, competitive with Fable 5 and Opus 4.8.
Terminal-Bench-LILT presents a multilingual coding benchmark with 300 authentic tasks in 10 languages, exposing gaps in AI models' handling of non-English issues like internationalization and cultural conventions.
A new paper claims that an ensemble of Qwen3.8-27B models achieves coding performance comparable to Fable-5 on LiveCodeBench, potentially at a significantly lower cost.
Qwen3.8-27B by Alibaba Qwen achieves 9th place on the Code Arena benchmark with 1595 points, outperforming larger models like Gemma 4-31B and reshaping the Pareto Frontier in coding performance.
The paper proposes Schrödinger Repo, an evaluation framework for coding agents that dynamically instantiates repositories to address data leakage in benchmarks, showing that current LLMs may depend on memorized cues.
GLM 5.3 benchmark results on SlopCodeBench show it scoring 47.1% on one subset and tying with Fable 5 and GPT-5.6 Sol on another, demonstrating that all models struggle with the unsaturated coding benchmark.
ProgramBench Vetted is a benchmark with 50 tasks that evaluate AI agents on reconstructing programs from runnable binaries, designed to study long-context coordination and enable dense reward for reinforcement learning.
The article compares three AI models—Qwen 3.8 27B, GPT-5.6 Terra, and Grok 4.6—on their ability to build a Three.js fragrance launch site, detailing their implementation strengths and potential issues.
Ramp, a spend management platform, launched its own AI research lab called Ramp Labs a year ago. The lab has worked on projects like a production-focused coding benchmark 'Ramp SWE-Bench', integrating Claude Code into RollerCoaster Tycoon, and a mechanistic interpretability playground.
Alibaba officially released Qwen3.8, with 2.4 trillion parameters and 1M context. Coding and office capabilities have been significantly improved, entering the global top tier. The author conducted a hardcore real-world test using a single-file HTML N-body simulation.
DeepSeek-V4-Flash-High tops the Frontend Code Arena with a 1586 score, offering the best performance-per-dollar at $0.14/$0.28 per MToken.
A Reddit user compares Inkling-Small-276B-12B and Qwen3.6-27B on a complex coding task, finding that Qwen plans and reviews its code methodically while Inkling produces hacky code after lengthy reasoning.
EvoCode-Bench is a multi-turn coding benchmark with 26 tasks across 5 domains, designed to evaluate AI agents on evolving specifications and cumulative testing in a persistent workspace, revealing that single-turn scores dramatically overstate reliability.
Poolside releases Laguna S 2.1, a 118B MoE model with 8B activated parameters per token, optimized for agentic coding. It claims to outperform DeepSeek V4 Pro while being cheaper than DeepSeek V4 Flash, with a 1M context window and open-source license.
Google announced Gemini 3.6 Flash with improved coding efficiency and lower token costs, alongside Gemini 3.5 Flash Lite and a cybersecurity-focused model, while still developing Gemini 3.5 Pro and hinting at Gemini 4.
Kimi K3 achieves top ranking on the Frontend Code Arena benchmark, demonstrating strong coding capabilities.
A detailed comparison of twelve AI models, including GPT-5.6, Grok 4.5, Claude, and open-weight models, tasked with building four different applications across multiple attempts, with all artifacts published for independent evaluation.