GPT-5.6 Luna vs GPT-6 Astra: is a $1.20 model good enough for code review?

Reddit r/artificial News

Summary

Benchmarked GPT-5.6 Luna vs GPT-6 Astra across 50 real PRs, showing Astra found 92 bugs vs Luna's 69, with Luna catching 75% of bugs at 3.6% of the cost. Includes detailed evaluation breakdown and upcoming comparison with Fable 5.1.

We benchmarked GPT-5.6 Luna vs GPT-6 Astra across 50 real PRs from Cal, Sentry, Discourse, Keycloak and Grafana. Astra found 92 confirmed bugs vs 69 for Luna, while Luna caught 75% of the bugs at just 3.6% of the cost. We also added the full eval breakdown this time, including cost, avg output tokens, latency, precision and bug classes like data/logic, security and concurrency. We’re doing Astra vs Fable 5.1 next, so would appreciate feedback on the evaluation before we run the next one. Dropping the link in the comments if anyone wants to check it out. https://preview.redd.it/xvin68lg9iph1.png?width=679&format=png&auto=webp&s=452996c144127d7b75cf1dd871d49376c1bcac0d
Original Article

Similar Articles

Gpt 6 astra benchmarks

Reddit r/singularity

This article covers the benchmarks for OpenAI's GPT-6 model, evaluating its performance using the Astra benchmark system.

GPT‑6 Astra

Simon Willison's Blog

OpenAI releases GPT-6 Astra, which excels in security tasks and long context handling, achieving 99.9% on ARC-AGI 3, though it still trails Claude Fable on some benchmarks.