arc-agi

Tag

Cards List
#arc-agi

Kepler: Auditable World Models for ARC-AGI-3

arXiv cs.AI ↗ · 18h ago Cached

Kepler is an open-source harness for ARC-AGI-3 that represents hypotheses as executable world models and validates them via retrospective transition and prediction checks, achieving a server-verified 100.00 RHAE on all 25 public games under a frozen Claude Opus 5 configuration. The paper also reports three evaluation failures — source-code leakage, harness reconstruction, and autonomous repair masking a broken planner — and argues that public-set scores alone have limited discriminative value, motivating first-attempt, cost-conditioned, and verification-aware reporting.

0 favorites 0 likes
#arc-agi

GPT-6 Sol & Astra dominate the ARC-AGI-3 leaderboard

Reddit r/singularity ↗ · 2d ago

GPT-6 Sol and Astra have taken the top spots on the ARC-AGI-3 leaderboard, marking a notable advance in abstract reasoning benchmarks for frontier AI models.

0 favorites 0 likes
#arc-agi

Reasoning with Neural Cellular Automata

arXiv cs.LG ↗ · 2d ago Cached

Google researchers show that Neural Cellular Automata with strictly local connectivity and asynchronous updates can solve complex visual reasoning tasks such as large mazes, Sudoku, and ARC-AGI-1, generalize out-of-distribution, and robustly recover from damage.

0 favorites 0 likes
#arc-agi

@ayhozade: Can neural networks learn to think by denoising? Excited to share Thinking with Looped Flows! We train recurrent reason…

X AI KOLs Timeline ↗ · 2026-09-11

A new approach called 'Thinking with Looped Flows' trains recurrent reasoning with local denoising objectives, achieving state-of-the-art performance on ARC-AGI benchmarks among looped models.

0 favorites 0 likes
#arc-agi

@ycombinator: Harnesses often get dismissed as just scaffolding, just prompt engineering, and not real research. But that couldn't be…

X AI KOLs Timeline ↗ · 2026-09-07 Cached

Y Combinator discusses the importance of harnesses in AI, highlighting their role in improving model performance, self-improving agents, and real-world applications such as personal AI and work automation.

0 favorites 0 likes
#arc-agi

Astra WITHOUT CoT gets 97% on ARC-AGI-3 and 86% on ARC-AGI-1

Reddit r/singularity ↗ · 2026-09-04

Astra achieves 97% on ARC-AGI-3 and 86% on ARC-AGI-1 without using Chain-of-Thought, highlighting a major advancement in AI reasoning capabilities.

0 favorites 0 likes
#arc-agi

GPT-6 is released [N]

Reddit r/MachineLearning ↗ · 2026-09-04

GPT-6 has been released by OpenAI, demonstrating about 60% accuracy on the ARC-AGI-3 benchmark without additional harnesses.

0 favorites 0 likes
#arc-agi

@gdb: arc-agi-3 is now saturated

X AI KOLs Timeline ↗ · 2026-09-03 Cached

GPT-6 Astra by OpenAI achieves state-of-the-art performance on the ARC-AGI-3 benchmark, scoring 63% and surpassing human performance on 96% of levels, demonstrating advanced symbolic modeling capabilities.

0 favorites 0 likes
#arc-agi

The prevalent problem of misleading benchmark reporting (re: Astra)

Reddit r/singularity ↗ · 2026-09-03

OpenAI's benchmark reporting for Astra on ARC-AGI-3 is misleading due to using different harnesses, and the performance gap is less dramatic under standard conditions.

0 favorites 0 likes
#arc-agi

@aimalysheva: meet @mostik_ai! what happens when you put 12 PhDs in one room for four months? first place on the ARC-AGI leaderboard,…

X AI KOLs Following ↗ · 2026-09-02 Cached

Russian startup Mostik has developed a method for AI models to communicate in latent space, enabling a small model to leverage reasoning from a frontier model without text, achieving 80% accuracy at 20x faster performance, and they are partnering with inference providers to promote open-weight adoption.

0 favorites 0 likes
#arc-agi

44% on ARC-AGI-1 in 67 cents

Hacker News Top ↗ · 2026-09-01 Cached

Trained a small transformer model from scratch to achieve 44% accuracy on the ARC-AGI-1 benchmark for only 67 cents, demonstrating improvements in speed, accuracy, and cost-effectiveness over previous methods.

0 favorites 0 likes
#arc-agi

Intelligence per dollar is the new scaling law: A tiny reasoning model breaks the existing cost-accuracy Pareto frontier on Arc-AGI 1

Reddit r/artificial ↗ · 2026-08-17

Chart Pathway's BDH-CQ, a 150M parameter reasoning model, achieves 29.5% on ARC-AGI-1 at a much lower cost per task compared to larger models like GPT-5.6 Luna, showcasing improved cost-accuracy trade-offs.

0 favorites 0 likes
#arc-agi

A 150M param recurrent model scores 29.5% on ARC-AGI-1 at $0.0007 per task

Reddit r/LocalLLaMA ↗ · 2026-08-14 Cached

The article introduces BDH-CQ, a 150M parameter recurrent model that combines in-context learning with latent reasoning, achieving 29.5% on ARC-AGI-1 at a cost of $0.0007 per task, setting a new standard for cost efficiency.

0 favorites 0 likes
#arc-agi

Transformer co-author validates post-transformer cost efficiency breakthrough

Reddit r/artificial ↗ · 2026-08-14

A 150M-parameter non-transformer architecture achieves state-of-the-art cost-efficiency on ARC-AGI-1, validated by Transformer co-author Łukasz Kaiser, suggesting that recurrent latent reasoning can replace brute-force scaling.

0 favorites 0 likes
#arc-agi

Did Pathway just reveal the architecture breakthrough Andrew Curran predicted? Its 150M model sets a new ARC-AGI-1 cost-efficiency frontier

Reddit r/singularity ↗ · 2026-08-11

Pathway's 150M-parameter BDH-CQ model achieves 29.5% on ARC-AGI-1 at a record-low cost of $0.0007 per task, using recurrent memory and latent reasoning instead of long token chains. The architecture may be the breakthrough Andrew Curran teased, with OpenAI researcher Lukasz Kaiser as an investor and adviser.

0 favorites 0 likes
#arc-agi

BDH-CQ: In-Context Learning with Recurrent Latent Reasoning

Hugging Face Daily Papers ↗ · 2026-08-10 Cached

This paper introduces BDH-CQ, a 150M-parameter reasoning model that combines in-context learning with recurrent latent reasoning, achieving 29.5% pass@2 on ARC-AGI-1 at very low inference cost and establishing a new cost-accuracy frontier.

0 favorites 0 likes
#arc-agi

We got 100% on ARC-AGI-3 ft09 with zero model calls. The failures are more interesting.

Reddit r/artificial ↗ · 2026-08-09

An experimental reasoning system at Orivael scored 100% on ARC-AGI-3 ft09 with zero model calls, revealing that its failures stem from incorrect environment representations rather than planning errors.

0 favorites 0 likes
#arc-agi

DeepSeek V4 Flash 0731

Hacker News Top ↗ · 2026-08-07 Cached

DeepSeek V4 Flash 0731 presents its results on the ARC-AGI benchmark, highlighting progress in abstract reasoning for AI models.

0 favorites 0 likes
#arc-agi

@reach_vb: luna-maxxing @ 80% lower costs!

X AI KOLs Timeline ↗ · 2026-08-07 Cached

ARC Prize re-tested OpenAI's GPT-5.6 Luna on ARC-AGI after an 80% price cut, confirming similar performance at a much lower cost per task.

0 favorites 0 likes
#arc-agi

Prime Agent - a new coding harness surpassing Codex/CC/PI

Reddit r/LocalLLaMA ↗ · 2026-08-05

Prime Agent is an open-source coding and research harness that outperforms proprietary harnesses, scoring 95.5% on ARC-AGI-3 and improving models across benchmarks.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback