Tag
Kepler is an open-source harness for ARC-AGI-3 that represents hypotheses as executable world models and validates them via retrospective transition and prediction checks, achieving a server-verified 100.00 RHAE on all 25 public games under a frozen Claude Opus 5 configuration. The paper also reports three evaluation failures — source-code leakage, harness reconstruction, and autonomous repair masking a broken planner — and argues that public-set scores alone have limited discriminative value, motivating first-attempt, cost-conditioned, and verification-aware reporting.
GPT-6 Sol and Astra have taken the top spots on the ARC-AGI-3 leaderboard, marking a notable advance in abstract reasoning benchmarks for frontier AI models.
Google researchers show that Neural Cellular Automata with strictly local connectivity and asynchronous updates can solve complex visual reasoning tasks such as large mazes, Sudoku, and ARC-AGI-1, generalize out-of-distribution, and robustly recover from damage.
A new approach called 'Thinking with Looped Flows' trains recurrent reasoning with local denoising objectives, achieving state-of-the-art performance on ARC-AGI benchmarks among looped models.
Y Combinator discusses the importance of harnesses in AI, highlighting their role in improving model performance, self-improving agents, and real-world applications such as personal AI and work automation.
Astra achieves 97% on ARC-AGI-3 and 86% on ARC-AGI-1 without using Chain-of-Thought, highlighting a major advancement in AI reasoning capabilities.
GPT-6 has been released by OpenAI, demonstrating about 60% accuracy on the ARC-AGI-3 benchmark without additional harnesses.
GPT-6 Astra by OpenAI achieves state-of-the-art performance on the ARC-AGI-3 benchmark, scoring 63% and surpassing human performance on 96% of levels, demonstrating advanced symbolic modeling capabilities.
OpenAI's benchmark reporting for Astra on ARC-AGI-3 is misleading due to using different harnesses, and the performance gap is less dramatic under standard conditions.
Russian startup Mostik has developed a method for AI models to communicate in latent space, enabling a small model to leverage reasoning from a frontier model without text, achieving 80% accuracy at 20x faster performance, and they are partnering with inference providers to promote open-weight adoption.
Trained a small transformer model from scratch to achieve 44% accuracy on the ARC-AGI-1 benchmark for only 67 cents, demonstrating improvements in speed, accuracy, and cost-effectiveness over previous methods.
Chart Pathway's BDH-CQ, a 150M parameter reasoning model, achieves 29.5% on ARC-AGI-1 at a much lower cost per task compared to larger models like GPT-5.6 Luna, showcasing improved cost-accuracy trade-offs.
The article introduces BDH-CQ, a 150M parameter recurrent model that combines in-context learning with latent reasoning, achieving 29.5% on ARC-AGI-1 at a cost of $0.0007 per task, setting a new standard for cost efficiency.
A 150M-parameter non-transformer architecture achieves state-of-the-art cost-efficiency on ARC-AGI-1, validated by Transformer co-author Łukasz Kaiser, suggesting that recurrent latent reasoning can replace brute-force scaling.
Pathway's 150M-parameter BDH-CQ model achieves 29.5% on ARC-AGI-1 at a record-low cost of $0.0007 per task, using recurrent memory and latent reasoning instead of long token chains. The architecture may be the breakthrough Andrew Curran teased, with OpenAI researcher Lukasz Kaiser as an investor and adviser.
This paper introduces BDH-CQ, a 150M-parameter reasoning model that combines in-context learning with recurrent latent reasoning, achieving 29.5% pass@2 on ARC-AGI-1 at very low inference cost and establishing a new cost-accuracy frontier.
An experimental reasoning system at Orivael scored 100% on ARC-AGI-3 ft09 with zero model calls, revealing that its failures stem from incorrect environment representations rather than planning errors.
DeepSeek V4 Flash 0731 presents its results on the ARC-AGI benchmark, highlighting progress in abstract reasoning for AI models.
ARC Prize re-tested OpenAI's GPT-5.6 Luna on ARC-AGI after an 80% price cut, confirming similar performance at a much lower cost per task.
Prime Agent is an open-source coding and research harness that outperforms proprietary harnesses, scoring 95.5% on ARC-AGI-3 and improving models across benchmarks.