Tag
The article introduces BDH-CQ, a 150M parameter recurrent model that combines in-context learning with latent reasoning, achieving 29.5% on ARC-AGI-1 at a cost of $0.0007 per task, setting a new standard for cost efficiency.
A 150M-parameter non-transformer architecture achieves state-of-the-art cost-efficiency on ARC-AGI-1, validated by Transformer co-author Łukasz Kaiser, suggesting that recurrent latent reasoning can replace brute-force scaling.
Pathway's 150M-parameter BDH-CQ model achieves 29.5% on ARC-AGI-1 at a record-low cost of $0.0007 per task, using recurrent memory and latent reasoning instead of long token chains. The architecture may be the breakthrough Andrew Curran teased, with OpenAI researcher Lukasz Kaiser as an investor and adviser.
This paper introduces BDH-CQ, a 150M-parameter reasoning model that combines in-context learning with recurrent latent reasoning, achieving 29.5% pass@2 on ARC-AGI-1 at very low inference cost and establishing a new cost-accuracy frontier.
An experimental reasoning system at Orivael scored 100% on ARC-AGI-3 ft09 with zero model calls, revealing that its failures stem from incorrect environment representations rather than planning errors.
DeepSeek V4 Flash 0731 presents its results on the ARC-AGI benchmark, highlighting progress in abstract reasoning for AI models.
ARC Prize re-tested OpenAI's GPT-5.6 Luna on ARC-AGI after an 80% price cut, confirming similar performance at a much lower cost per task.
Prime Agent is an open-source coding and research harness that outperforms proprietary harnesses, scoring 95.5% on ARC-AGI-3 and improving models across benchmarks.
Linus Ekenstam highlights Prime Intellect's release of Prime Agent, a self-improving harness for coding and long-running autonomous tasks, reportedly scoring 95.5% on ARC-AGI-3, above the human baseline.
An analysis comparing Claude Opus 5 High and GPT 5.6 Sol Max on an ARC-AGI-3 puzzle shows Opus winning by preserving detailed state in visible output, while Sol relies on discarded hidden reasoning.
Discusses how a 25% difference on ARC-AGI was due to harness setup, showing GPT-5.6 Sol scoring 38% with proper evaluation, and critiques naive benchmark reporting in the industry.
A critique arguing that ARC-AGI 3 unfairly disables an agent's ability to maintain context across actions, making it an dishonest measure of general intelligence. It notes that allowing compaction triples scores while using far fewer tokens, and that real-world agents work that way.
Seed IQ demonstrates advanced real-time perception, reasoning, and adaptation in dynamic environments by navigating Doom II, potentially surpassing static benchmarks like ARC AGI 3.
A discussion about whether the general consensus on ARC AGI 3 is that it cannot be 'benchmaxxed' (optimized for the benchmark), seeking opinions on the topic.
Opus 5 achieved a high score on the ARC AGI benchmark, indicating advanced reasoning capabilities.
The ARC-AGI leaderboard shows model performance on three versions of the benchmark, measuring fluid intelligence and efficient adaptation, with trend lines for reasoning systems and raw LLMs.
Discusses a potential vulnerability in the ARC AGI 3 benchmark where the Opus model could be gamed if it functions as a loop rather than a pure model.
Opus 5 achieves 30.2% on the ARC-AGI3 benchmark, marking a notable performance improvement.
Sakana AI introduces AB-MCTS, a new inference-time scaling algorithm that enables multiple frontier AI models to cooperate, significantly improving performance on the ARC-AGI-2 benchmark.
This paper investigates whether coding agents require executable world models, simplification, and verification to solve the ARC-AGI-3 benchmark, contributing to research on AGI and reasoning.