coding-benchmark

Tag

Cards List
#coding-benchmark

A 35B model beat a 120B one on my coding agent, 95% vs 53%. Build your own benchmark.

Reddit r/AI_Agents ↗ · 2026-09-26

A developer built a custom benchmark for coding AI agents and found that a 35B-parameter model outperformed a 120B-parameter one when the harness was optimized, highlighting the importance of tailored evaluation over generic specifications.

0 favorites 0 likes
#coding-benchmark

Qwen3.8-Max-0902 Beats Claude Opus 5 on Coding

Reddit r/singularity ↗ · 2026-09-02 Cached

Qwen3.8-Max-0902 debuts at #1 in Code Arena: WebDev with 1691 points, surpassing Claude Opus 5 by 3 points, and claims the highest-scoring position on the Pareto frontier at $5/MToken.

0 favorites 0 likes
#coding-benchmark

@Suhail: Excited to see benchmarks headed in this direction and getting better!

X AI KOLs Timeline ↗ · 2026-09-02 Cached

Proximal has released FrontierSWE v2, an updated ultra-long horizon coding benchmark with expanded tasks and improved methodology, highlighting large performance gaps where Claude Fable 5.1 leads.

0 favorites 0 likes
#coding-benchmark

@Yuchenj_UW: GLM-5.3 at 310 tok/s! Databricks inference is #1 in both speed and latency, again. On our internal Databricks coding be…

X AI KOLs Following ↗ · 2026-09-02 Cached

GLM-5.3 achieves 310 tokens per second on Databricks inference, leading in both speed and latency, and is the strongest open-source model for coding, competitive with Fable 5 and Opus 4.8.

0 favorites 0 likes
#coding-benchmark

Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culture

arXiv cs.CL ↗ · 2026-09-01 Cached

Terminal-Bench-LILT presents a multilingual coding benchmark with 300 authentic tasks in 10 languages, exposing gaps in AI models' handling of non-English issues like internationalization and cultural conventions.

0 favorites 0 likes
#coding-benchmark

Do you think a few Qwen3.8-27B models working together could score as well as Fable-5 on LiveCodeBench Hard?

Reddit r/LocalLLaMA ↗ · 2026-08-28

A new paper claims that an ensemble of Qwen3.8-27B models achieves coding performance comparable to Fable-5 on LiveCodeBench, potentially at a significantly lower cost.

0 favorites 0 likes
#coding-benchmark

Qwen 3.8 27B in 9th position on code arena. Gemma 4 31B is 80th.

Reddit r/LocalLLaMA ↗ · 2026-08-24 Cached

Qwen3.8-27B by Alibaba Qwen achieves 9th place on the Code Arena benchmark with 1595 points, outperforming larger models like Gemma 4-31B and reshaping the Pareto Frontier in coding performance.

0 favorites 0 likes
#coding-benchmark

Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It?

Hugging Face Daily Papers ↗ · 2026-08-21 Cached

The paper proposes Schrödinger Repo, an evaluation framework for coding agents that dynamically instantiates repositories to address data leakage in benchmarks, showing that current LLMs may depend on memorized cues.

0 favorites 0 likes
#coding-benchmark

GLM 5.3 SlopCodeBench Results

Reddit r/LocalLLaMA ↗ · 2026-08-20

GLM 5.3 benchmark results on SlopCodeBench show it scoring 47.1% on one subset and tying with Fable 5 and GPT-5.6 Sol on another, demonstrating that all models struggle with the unsaturated coding benchmark.

0 favorites 0 likes
#coding-benchmark

ProgramBench Vetted: Reverse Engineering from a Runnable Binary

Hacker News Top ↗ · 2026-08-20 Cached

ProgramBench Vetted is a benchmark with 50 tasks that evaluate AI agents on reconstructing programs from runnable binaries, designed to study long-context coordination and enable dense reward for reinforcement learning.

0 favorites 0 likes
#coding-benchmark

Local Qwen 3.8 27B vs GPT‑5.6 Terra vs Grok 4.6

Reddit r/artificial ↗ · 2026-08-18

The article compares three AI models—Qwen 3.8 27B, GPT-5.6 Terra, and Grok 4.6—on their ability to build a Three.js fragrance launch site, detailing their implementation strengths and potential issues.

0 favorites 0 likes
#coding-benchmark

@Redpoint: Why did a spend management platform start its own AI research lab? @karimatiyeh describes @RampLabs as a “collection of…

X AI KOLs Following ↗ · 2026-08-07 Cached

Ramp, a spend management platform, launched its own AI research lab called Ramp Labs a year ago. The lab has worked on projects like a production-focused coding benchmark 'Ramp SWE-Bench', integrating Claude Code into RollerCoaster Tycoon, and a mechanistic interpretability playground.

0 favorites 0 likes
#coding-benchmark

@GoSailGlobal: Today, August 3, Alibaba officially released Qwen3.8. With 2.4 trillion parameters and 1M context, coding and office capabilities have improved dramatically in this version, pushing it into the global top tier overall. The preview version ran for two weeks, and today it officially graduated. I immediately put it through a hardcore real-world test—let me lay out the results and process. I came up with a pretty tough challenge: single-file HTML hand...

X AI KOLs Timeline ↗ · 2026-08-03 Cached

Alibaba officially released Qwen3.8, with 2.4 trillion parameters and 1M context. Coding and office capabilities have been significantly improved, entering the global top tier. The author conducted a hardcore real-world test using a single-file HTML N-body simulation.

0 favorites 0 likes
#coding-benchmark

@omarsar0: This DeepSeek-V4-Flash-High model is insanely good at front end. I ran some tests and had to double check if I had the …

X AI KOLs Following ↗ · 2026-08-01 Cached

DeepSeek-V4-Flash-High tops the Frontend Code Arena with a 1586 score, offering the best performance-per-dollar at $0.14/$0.28 per MToken.

0 favorites 0 likes
#coding-benchmark

Inkling-Small-276B-12B, effort "max" VS Qwen3.6-27B

Reddit r/LocalLLaMA ↗ · 2026-07-30

A Reddit user compares Inkling-Small-276B-12B and Qwen3.6-27B on a complex coding task, finding that Qwen plans and reviews its code methodically while Inkling produces hacky code after lengthy reasoning.

0 favorites 0 likes
#coding-benchmark

@_philschmid: https://x.com/_philschmid/status/2081744861829414977

X AI KOLs Timeline ↗ · 2026-07-27 Cached

EvoCode-Bench is a multi-turn coding benchmark with 26 tasks across 5 domains, designed to evaluate AI agents on evolving specifications and cumulative testing in a persistent workspace, revealing that single-turn scores dramatically overstate reliability.

0 favorites 0 likes
#coding-benchmark

Laguna S 2.1 Released: Cheaper than Deepseek v4 Flash, Better than V4 Pro

Reddit r/LocalLLaMA ↗ · 2026-07-21 Cached

Poolside releases Laguna S 2.1, a 118B MoE model with 8B activated parameters per token, optimized for agentic coding. It claims to outperform DeepSeek V4 Pro while being cheaper than DeepSeek V4 Flash, with a 1M context window and open-source license.

0 favorites 0 likes
#coding-benchmark

Google announces Gemini 3.6 Flash and cybersecurity AI, teases 3.5 Pro and Gemini 4

Ars Technica ↗ · 2026-07-21 Cached

Google announced Gemini 3.6 Flash with improved coding efficiency and lower token costs, alongside Gemini 3.5 Flash Lite and a cybersecurity-focused model, while still developing Gemini 3.5 Pro and hinting at Gemini 4.

0 favorites 0 likes
#coding-benchmark

Kimi K3 tops Frontend Code Arena

Reddit r/singularity ↗ · 2026-07-16

Kimi K3 achieves top ranking on the Frontend Code Arena benchmark, demonstrating strong coding capabilities.

0 favorites 0 likes
#coding-benchmark

GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

Hacker News Top ↗ · 2026-07-10 Cached

A detailed comparison of twelve AI models, including GPT-5.6, Grok 4.5, Claude, and open-weight models, tasked with building four different applications across multiple attempts, with all artifacts published for independent evaluation.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback