Tag
Ramp, a spend management platform, launched its own AI research lab called Ramp Labs a year ago. The lab has worked on projects like a production-focused coding benchmark 'Ramp SWE-Bench', integrating Claude Code into RollerCoaster Tycoon, and a mechanistic interpretability playground.
Alibaba officially released Qwen3.8, with 2.4 trillion parameters and 1M context. Coding and office capabilities have been significantly improved, entering the global top tier. The author conducted a hardcore real-world test using a single-file HTML N-body simulation.
DeepSeek-V4-Flash-High tops the Frontend Code Arena with a 1586 score, offering the best performance-per-dollar at $0.14/$0.28 per MToken.
A Reddit user compares Inkling-Small-276B-12B and Qwen3.6-27B on a complex coding task, finding that Qwen plans and reviews its code methodically while Inkling produces hacky code after lengthy reasoning.
EvoCode-Bench is a multi-turn coding benchmark with 26 tasks across 5 domains, designed to evaluate AI agents on evolving specifications and cumulative testing in a persistent workspace, revealing that single-turn scores dramatically overstate reliability.
Poolside releases Laguna S 2.1, a 118B MoE model with 8B activated parameters per token, optimized for agentic coding. It claims to outperform DeepSeek V4 Pro while being cheaper than DeepSeek V4 Flash, with a 1M context window and open-source license.
Google announced Gemini 3.6 Flash with improved coding efficiency and lower token costs, alongside Gemini 3.5 Flash Lite and a cybersecurity-focused model, while still developing Gemini 3.5 Pro and hinting at Gemini 4.
Kimi K3 achieves top ranking on the Frontend Code Arena benchmark, demonstrating strong coding capabilities.
A detailed comparison of twelve AI models, including GPT-5.6, Grok 4.5, Claude, and open-weight models, tasked with building four different applications across multiple attempts, with all artifacts published for independent evaluation.
Grok 4.5 with Grok Build achieved #1 on the SWE-Atlas-QnA benchmark with a score of 84, matching GPT-5.6 Codex and outperforming other coding setups.
GPT5.6 Sol Ultra achieves 91.9% on TerminalBench coding benchmark, suggesting coding tasks are approaching solved.
This article benchmarks Grok 4.5, GPT-5.5, Claude Opus 4.8, and Claude Fable 5 by having each model build three interactive apps (3D Rubik's Cube, particle gravity sandbox, Breakout game) from a single prompt, comparing their one-shot coding capabilities.
The May 2025 Sonnet beats Sonnet 5 on LiveBench's general coding score but loses by 27 points on agentic coding, highlighting differences in benchmark performance.
Cognition launches SWE-1.7, a highly capable AI model for agentic software engineering that achieves frontier-level performance at reduced cost, with improvements in RL training, multi-cluster infrastructure, data curation, and self-compaction for long tasks.
A user reports that Mimo v2.5 outperforms DeepSeek v4 Flash in coding tasks based on benchmarks like Codex, Oh My Pi, Hermes, and Terminal Bench v2.0, though both models are similar overall.
Ornith-1.0-9B is a new 9B parameter AI model optimized for 8-12GB GPUs, achieving strong performance on agentic coding benchmarks, matching or surpassing models 2-3x its size.
Devin Desktop now supports Kimi K2.7 and GLM 5.2 models, offering free trials until July 5 for Pro/Max/Teams users.
GLM 5.2 is a new open-weights model from Z.ai, compared against Claude Opus in a 3D game coding task. Opus performed faster and cleaner, but GLM 5.2 offers compelling cost and accessibility advantages.
Step 3.7 Flash, an open-weights model with a 256k context window, is available free in Cline for a month, claiming to outperform Gemini and DeepSeek flash models and approach frontier performance on SWE Bench.
The writer shares their experience with Nex-N2 Pro, originally mistaken as Rio-3.5, and finds it performs exceptionally well on coding benchmarks without hallucination, rivaling GPT-5.x on their Mac setup.