Astra scores the highest on Blueprint-bench2
Summary
Astra achieved the highest score on Blueprint-Bench 2, a benchmark testing AI agents' ability to convert apartment photos into accurate 2D floor plans using spatial reasoning and cross-apartment learning.
Similar Articles
Astra leads in IKEA furniture assembly
Epoch AI created the Furniture Assembly Benchmark (FAB) to test AI models' visual reasoning in spotting mistakes in IKEA assembly photos. OpenAI's GPT-6 Astra leads with 80% accuracy, showing major improvements over previous models.
@BenjaminDEKR: Astra understands spatial / 3D relationships better than other leading models. This is why it's so good at CAD, models,…
Astra reportedly surpasses other leading models in spatial and 3D understanding, achieving first place on the VoxelBench benchmark with an Elo rating exceeding 2600 and a lead of over 300 points.
OpenAI's Astra scored 62.7% and 99.9% on the same benchmark, and I still don't fully know which one to trust
The article examines inconsistencies in benchmark scores for OpenAI's Astra model on the ARC Prize, highlighting discrepancies from different testing harnesses and changes to OpenAI's launch page, which prompts concerns about evaluation accuracy.
Not only does Astra saturate ARC-AGI-3, it does so using fewer moves than the average human
GPT-6 Astra achieves state-of-the-art scores on the ARC-AGI-3 benchmark, scoring 99.9% with a Provider Adapter harness and demonstrating fewer actions than human testers. The model exhibits the ability to convert unfamiliar environments into compact symbolic world models for efficient planning.
@reach_vb: Astra is SoTA on MazeBench by a huge margin:
Astra achieves state-of-the-art performance on MazeBench, significantly outperforming GPT-6 in a 3D open world spatial reasoning evaluation.