@reach_vb: Astra is SoTA on MazeBench by a huge margin:
Summary
Astra achieves state-of-the-art performance on MazeBench, significantly outperforming GPT-6 in a 3D open world spatial reasoning evaluation.
View Cached Full Text
Cached at: 09/08/26, 05:17 AM
Astra is SoTA on MazeBench by a huge margin: https://t.co/YEBrBMfcRh
💺 (@patience_cave): MazeBench vs GPT-6 Astra
Astra spent 60+ hours in this 3D open world spatial reasoning eval.
Final score: 14%
Similar Articles
Not only does Astra saturate ARC-AGI-3, it does so using fewer moves than the average human
GPT-6 Astra achieves state-of-the-art scores on the ARC-AGI-3 benchmark, scoring 99.9% with a Provider Adapter harness and demonstrating fewer actions than human testers. The model exhibits the ability to convert unfamiliar environments into compact symbolic world models for efficient planning.
@j_dekoninck: GPT-6-Astra takes first place on MathArena, with a massive 90% expected performance! It is also very token efficient: o…
GPT-6-Astra takes first place on MathArena with a 90% expected performance and high token efficiency.
GPT-6 Astra Pro creates a maze on MineBench
GPT-6 Astra Pro demonstrates the capability to create a polished maze with a correct solution on MineBench, highlighting advanced AI problem-solving skills.
GPT‑6 Astra
OpenAI releases GPT-6 Astra, which excels in security tasks and long context handling, achieving 99.9% on ARC-AGI 3, though it still trails Claude Fable on some benchmarks.
Gpt 6 astra benchmarks
This article covers the benchmarks for OpenAI's GPT-6 model, evaluating its performance using the Astra benchmark system.