arc-agi

Tag

Cards List
#arc-agi

A 150M param recurrent model scores 29.5% on ARC-AGI-1 at $0.0007 per task

Reddit r/LocalLLaMA · 2d ago Cached

The article introduces BDH-CQ, a 150M parameter recurrent model that combines in-context learning with latent reasoning, achieving 29.5% on ARC-AGI-1 at a cost of $0.0007 per task, setting a new standard for cost efficiency.

0 favorites 0 likes
#arc-agi

Transformer co-author validates post-transformer cost efficiency breakthrough

Reddit r/artificial · 2d ago

A 150M-parameter non-transformer architecture achieves state-of-the-art cost-efficiency on ARC-AGI-1, validated by Transformer co-author Łukasz Kaiser, suggesting that recurrent latent reasoning can replace brute-force scaling.

0 favorites 0 likes
#arc-agi

Did Pathway just reveal the architecture breakthrough Andrew Curran predicted? Its 150M model sets a new ARC-AGI-1 cost-efficiency frontier

Reddit r/singularity · 5d ago

Pathway's 150M-parameter BDH-CQ model achieves 29.5% on ARC-AGI-1 at a record-low cost of $0.0007 per task, using recurrent memory and latent reasoning instead of long token chains. The architecture may be the breakthrough Andrew Curran teased, with OpenAI researcher Lukasz Kaiser as an investor and adviser.

0 favorites 0 likes
#arc-agi

BDH-CQ: In-Context Learning with Recurrent Latent Reasoning

Hugging Face Daily Papers · 2026-08-10 Cached

This paper introduces BDH-CQ, a 150M-parameter reasoning model that combines in-context learning with recurrent latent reasoning, achieving 29.5% pass@2 on ARC-AGI-1 at very low inference cost and establishing a new cost-accuracy frontier.

0 favorites 0 likes
#arc-agi

We got 100% on ARC-AGI-3 ft09 with zero model calls. The failures are more interesting.

Reddit r/artificial · 2026-08-09

An experimental reasoning system at Orivael scored 100% on ARC-AGI-3 ft09 with zero model calls, revealing that its failures stem from incorrect environment representations rather than planning errors.

0 favorites 0 likes
#arc-agi

DeepSeek V4 Flash 0731

Hacker News Top · 2026-08-07 Cached

DeepSeek V4 Flash 0731 presents its results on the ARC-AGI benchmark, highlighting progress in abstract reasoning for AI models.

0 favorites 0 likes
#arc-agi

@reach_vb: luna-maxxing @ 80% lower costs!

X AI KOLs Timeline · 2026-08-07 Cached

ARC Prize re-tested OpenAI's GPT-5.6 Luna on ARC-AGI after an 80% price cut, confirming similar performance at a much lower cost per task.

0 favorites 0 likes
#arc-agi

Prime Agent - a new coding harness surpassing Codex/CC/PI

Reddit r/LocalLLaMA · 2026-08-05

Prime Agent is an open-source coding and research harness that outperforms proprietary harnesses, scoring 95.5% on ARC-AGI-3 and improving models across benchmarks.

0 favorites 0 likes
#arc-agi

@LinusEkenstam: Just before bed time. Let me sleep. plz 95.5% on ARC-AGI-3 (’huge if true”)

X AI KOLs Timeline · 2026-08-05 Cached

Linus Ekenstam highlights Prime Intellect's release of Prime Agent, a self-improving harness for coding and long-running autonomous tasks, reportedly scoring 95.5% on ARC-AGI-3, above the human baseline.

0 favorites 0 likes
#arc-agi

I asked Sol Max to compare the output of Claude Opus 5 High and GPT 5.6 Sol Max on a specific puzzle on ARC-AGI-3 where Opus 5 had a 98.81% score and GPT 5.6 Sol Max had a 21.42% score

Reddit r/singularity · 2026-07-31

An analysis comparing Claude Opus 5 High and GPT 5.6 Sol Max on an ARC-AGI-3 puzzle shows Opus winning by preserving detailed state in visible output, while Sol relies on discarded hidden reasoning.

0 favorites 0 likes
#arc-agi

25% difference on a benchmark – just a mistake on how you run it😬

Reddit r/AI_Agents · 2026-07-30

Discusses how a 25% difference on ARC-AGI was due to harness setup, showing GPT-5.6 Sol scoring 38% with proper evaluation, and critiques naive benchmark reporting in the industry.

0 favorites 0 likes
#arc-agi

ARC-AGI 3 is not an honest measure of AGI

Reddit r/singularity · 2026-07-30

A critique arguing that ARC-AGI 3 unfairly disables an agent's ability to maintain context across actions, making it an dishonest measure of general intelligence. It notes that allowing compaction triples scores while using far fewer tokens, and that real-world agents work that way.

0 favorites 0 likes
#arc-agi

Seed IQ: Beyond ARC AGI 3? Watch It Navigate Doom II.

Reddit r/ArtificialInteligence · 2026-07-27

Seed IQ demonstrates advanced real-time perception, reasoning, and adaptation in dynamic environments by navigating Doom II, potentially surpassing static benchmarks like ARC AGI 3.

0 favorites 0 likes
#arc-agi

Was the general consensus of ARC AGI 3 was that it can't be benchmaxxed? What's your opinion?

Reddit r/singularity · 2026-07-25

A discussion about whether the general consensus on ARC AGI 3 is that it cannot be 'benchmaxxed' (optimized for the benchmark), seeking opinions on the topic.

0 favorites 0 likes
#arc-agi

Opus 5 ARC AGI score was benchmaxxed

Reddit r/singularity · 2026-07-25

Opus 5 achieved a high score on the ARC AGI benchmark, indicating advanced reasoning capabilities.

0 favorites 0 likes
#arc-agi

ARC-AGI Leaderboard

Hacker News Top · 2026-07-25 Cached

The ARC-AGI leaderboard shows model performance on three versions of the benchmark, measuring fluid intelligence and efficient adaptation, with trend lines for reasoning systems and raw LLMs.

0 favorites 0 likes
#arc-agi

ARC AGI 3 could be gamed if Opus is a loop and not a pure model

Reddit r/singularity · 2026-07-24

Discusses a potential vulnerability in the ARC AGI 3 benchmark where the Opus model could be gamed if it functions as a loop rather than a pure model.

0 favorites 0 likes
#arc-agi

Opus 5 benchmarks (30.2% on ARC-AGI3!!!)

Reddit r/singularity · 2026-07-24

Opus 5 achieves 30.2% on the ARC-AGI3 benchmark, marking a notable performance improvement.

0 favorites 0 likes
#arc-agi

Announcing Fugu-Ultra v1.1 🐡 (1 minute read)

TLDR AI · 2026-07-24 Cached

Sakana AI introduces AB-MCTS, a new inference-time scaling algorithm that enables multiple frontier AI models to cooperate, significantly improving performance on the ARC-AGI-2 benchmark.

0 favorites 0 likes
#arc-agi

Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?

arXiv cs.AI · 2026-07-20 Cached

This paper investigates whether coding agents require executable world models, simplification, and verification to solve the ARC-AGI-3 benchmark, contributing to research on AGI and reasoning.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback