coding-benchmark

Tag

Cards List
#coding-benchmark

@Redpoint: Why did a spend management platform start its own AI research lab? @karimatiyeh describes @RampLabs as a “collection of…

X AI KOLs Following · 2026-08-07 Cached

Ramp, a spend management platform, launched its own AI research lab called Ramp Labs a year ago. The lab has worked on projects like a production-focused coding benchmark 'Ramp SWE-Bench', integrating Claude Code into RollerCoaster Tycoon, and a mechanistic interpretability playground.

0 favorites 0 likes
#coding-benchmark

@GoSailGlobal: Today, August 3, Alibaba officially released Qwen3.8. With 2.4 trillion parameters and 1M context, coding and office capabilities have improved dramatically in this version, pushing it into the global top tier overall. The preview version ran for two weeks, and today it officially graduated. I immediately put it through a hardcore real-world test—let me lay out the results and process. I came up with a pretty tough challenge: single-file HTML hand...

X AI KOLs Timeline · 2026-08-03 Cached

Alibaba officially released Qwen3.8, with 2.4 trillion parameters and 1M context. Coding and office capabilities have been significantly improved, entering the global top tier. The author conducted a hardcore real-world test using a single-file HTML N-body simulation.

0 favorites 0 likes
#coding-benchmark

@omarsar0: This DeepSeek-V4-Flash-High model is insanely good at front end. I ran some tests and had to double check if I had the …

X AI KOLs Following · 2026-08-01 Cached

DeepSeek-V4-Flash-High tops the Frontend Code Arena with a 1586 score, offering the best performance-per-dollar at $0.14/$0.28 per MToken.

0 favorites 0 likes
#coding-benchmark

Inkling-Small-276B-12B, effort "max" VS Qwen3.6-27B

Reddit r/LocalLLaMA · 2026-07-30

A Reddit user compares Inkling-Small-276B-12B and Qwen3.6-27B on a complex coding task, finding that Qwen plans and reviews its code methodically while Inkling produces hacky code after lengthy reasoning.

0 favorites 0 likes
#coding-benchmark

@_philschmid: https://x.com/_philschmid/status/2081744861829414977

X AI KOLs Timeline · 2026-07-27 Cached

EvoCode-Bench is a multi-turn coding benchmark with 26 tasks across 5 domains, designed to evaluate AI agents on evolving specifications and cumulative testing in a persistent workspace, revealing that single-turn scores dramatically overstate reliability.

0 favorites 0 likes
#coding-benchmark

Laguna S 2.1 Released: Cheaper than Deepseek v4 Flash, Better than V4 Pro

Reddit r/LocalLLaMA · 2026-07-21 Cached

Poolside releases Laguna S 2.1, a 118B MoE model with 8B activated parameters per token, optimized for agentic coding. It claims to outperform DeepSeek V4 Pro while being cheaper than DeepSeek V4 Flash, with a 1M context window and open-source license.

0 favorites 0 likes
#coding-benchmark

Google announces Gemini 3.6 Flash and cybersecurity AI, teases 3.5 Pro and Gemini 4

Ars Technica · 2026-07-21 Cached

Google announced Gemini 3.6 Flash with improved coding efficiency and lower token costs, alongside Gemini 3.5 Flash Lite and a cybersecurity-focused model, while still developing Gemini 3.5 Pro and hinting at Gemini 4.

0 favorites 0 likes
#coding-benchmark

Kimi K3 tops Frontend Code Arena

Reddit r/singularity · 2026-07-16

Kimi K3 achieves top ranking on the Frontend Code Arena benchmark, demonstrating strong coding capabilities.

0 favorites 0 likes
#coding-benchmark

GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps

Hacker News Top · 2026-07-10 Cached

A detailed comparison of twelve AI models, including GPT-5.6, Grok 4.5, Claude, and open-weight models, tasked with building four different applications across multiple attempts, with all artifacts published for independent evaluation.

0 favorites 0 likes
#coding-benchmark

@elonmusk: Grok Build

X AI KOLs Timeline · 2026-07-10 Cached

Grok 4.5 with Grok Build achieved #1 on the SWE-Atlas-QnA benchmark with a score of 84, matching GPT-5.6 Codex and outperforming other coding setups.

0 favorites 0 likes
#coding-benchmark

@BenjaminDEKR: So it's just over? GPT5.6 Sol Ultra scores 91.9% on TerminalBench Coding is approaching solved, the same way arithmetic…

X AI KOLs Following · 2026-07-09 Cached

GPT5.6 Sol Ultra achieves 91.9% on TerminalBench coding benchmark, suggesting coding tasks are approaching solved.

0 favorites 0 likes
#coding-benchmark

We made Grok 4.5, GPT-5.5, and Claude build the same apps

Reddit r/singularity · 2026-07-09 Cached

This article benchmarks Grok 4.5, GPT-5.5, Claude Opus 4.8, and Claude Fable 5 by having each model build three interactive apps (3D Rubik's Cube, particle gravity sandbox, Breakout game) from a single prompt, comparing their one-shot coding capabilities.

0 favorites 0 likes
#coding-benchmark

The May 2025 Sonnet still beats Sonnet 5 on livebench's coding score. On agentic coding it lose to it by 27 points. So what's the difference?

Reddit r/AI_Agents · 2026-07-09

The May 2025 Sonnet beats Sonnet 5 on LiveBench's general coding score but loses by 27 points on agentic coding, highlighting differences in benchmark performance.

0 favorites 0 likes
#coding-benchmark

SWE-1.7: Frontier Intelligence at a Fraction of the Cost (22 minute read)

TLDR AI · 2026-07-09 Cached

Cognition launches SWE-1.7, a highly capable AI model for agentic software engineering that achieves frontier-level performance at reduced cost, with improvements in RL training, multi-cluster infrastructure, data curation, and self-compaction for long tasks.

0 favorites 0 likes
#coding-benchmark

any one else finds Mimo v2.5 better than deepseek v4 flash!?

Reddit r/LocalLLaMA · 2026-07-08

A user reports that Mimo v2.5 outperforms DeepSeek v4 Flash in coding tasks based on benchmarks like Codex, Oh My Pi, Hermes, and Terminal Bench v2.0, though both models are similar overall.

0 favorites 0 likes
#coding-benchmark

@SlimTradeyBaby: Attention all 8-12GB GPU users! This new Ornith-1.0-9B is looking like it will be a serz player for smaller VRAM setups…

X AI KOLs Timeline · 2026-06-26 Cached

Ornith-1.0-9B is a new 9B parameter AI model optimized for 8-12GB GPUs, achieving strong performance on agentic coding benchmarks, matching or surpassing models 2-3x its size.

0 favorites 0 likes
#coding-benchmark

@cognition: Try Kimi K2.7 and GLM 5.2 for free in Devin Desktop and CLI

X AI KOLs Following · 2026-06-24 Cached

Devin Desktop now supports Kimi K2.7 and GLM 5.2 models, offering free trials until July 5 for Pro/Max/Teams users.

0 favorites 0 likes
#coding-benchmark

GLM 5.2 vs. Opus

Hacker News Top · 2026-06-22 Cached

GLM 5.2 is a new open-weights model from Z.ai, compared against Claude Opus in a 3D game coding task. Opus performed faster and cleaner, but GLM 5.2 offers compelling cost and accessibility advantages.

0 favorites 0 likes
#coding-benchmark

@cline: Step 3.7 Flash is free in Cline for the next month. It beats Gemini and DeepSeek flash models, and comes surprisingly c…

X AI KOLs Following · 2026-06-17 Cached

Step 3.7 Flash, an open-weights model with a 256k context window, is available free in Cline for a month, claiming to outperform Gemini and DeepSeek flash models and approach frontier performance on SWE Bench.

0 favorites 0 likes
#coding-benchmark

Nex-N2 Pro is the real deal

Reddit r/LocalLLaMA · 2026-06-16

The writer shares their experience with Nex-N2 Pro, originally mistaken as Rio-3.5, and finds it performs exceptionally well on coding benchmarks without hallucination, rivaling GPT-5.x on their Mac setup.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback