arc-agi

Tag

Cards List
#arc-agi

Schema (2 minute read)

TLDR AI · 2026-07-17

Schema is a harness that enables frontier AI models to achieve 99% on the ARC-AGI-3 benchmark by having them write executable programs to model game environments, test predictions, and plan.

0 favorites 0 likes
#arc-agi

@HavenFeng: Today, we’re introducing [schema]: a harness reaching 99% RHAE with Opus 4.8 + Fable 5 and 95.35% with GPT-5.6 Sol on A…

X AI KOLs Timeline · 2026-07-16 Cached

Introducing [schema], a harness that achieves 99% RHAE with Opus 4.8 + Fable 5 and 95.35% with GPT-5.6 Sol on ARC-AGI-3 Public set, designed to make an LLM think like a physicist.

0 favorites 0 likes
#arc-agi

GPT-5.6 Series (2 minute read)

TLDR AI · 2026-07-10 Cached

OpenAI releases the GPT-5.6 model series, with Sol being the standout model achieving 13.33% on ARC-AGI-3 public and the first to win a game, demonstrating improved ability to orient itself in unfamiliar environments.

0 favorites 0 likes
#arc-agi

ChatGPT 5.6 - ARC-AGI 3 score

Reddit r/singularity · 2026-07-09

ChatGPT 5.6 achieved a new score on the ARC-AGI benchmark, indicating progress toward general intelligence.

0 favorites 0 likes
#arc-agi

Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1

arXiv cs.AI · 2026-07-09 Cached

This paper presents cost-effective agent harnesses for ARC-AGI-1 that achieve strong performance using DeepSeek V3.2 without fine-tuning, via an Explorer-Definer Pipeline and a Reflective Orchestrator, achieving 67.25% pass@2 at low cost.

0 favorites 0 likes
#arc-agi

@alxfazio: if you’re into harness engineering, i strongly recommend looking into arc agi winning harnesses. they clearly illustrat…

X AI KOLs Timeline · 2026-07-03

A recommendation to examine winning harnesses from the ARC AGI challenge to understand effective first-principles design and avoid overfitting to benchmarks.

0 favorites 0 likes
#arc-agi

OPINE-World: Programmatic World Modeling with Ontology-error-Prioritized Interactive Exploration

arXiv cs.AI · 2026-07-03 Cached

OPINE-World introduces an LLM agent that learns an object-centric programmatic world model online through interaction, using ontology-error-prioritized exploration and cooperating hypothesis-test agents, achieving strong results on ARC-AGI-3.

0 favorites 0 likes
#arc-agi

@sethkarten: https://x.com/sethkarten/status/2072034978112889328

X AI KOLs Following · 2026-06-30 Cached

Continual Harness is a reset-free, self-improving agentic harness that achieves 20.54% on ARC-AGI-3 at a cost of $774 by storing memories, reusing skills, and refining its prompt, outperforming prior baselines like Hermes and OpenClaw with greater efficiency.

0 favorites 0 likes
#arc-agi

@GregKamradt: Alexia was the solo author on TRM, pushing the limits with recursion on a small 2 layer network TRM’s performance on AR…

X AI KOLs Timeline · 2026-06-30 Cached

Alexia Jolicoeur-Martineau was the solo author on TRM, pushing recursion limits on a small 2-layer network, achieving impressive performance on ARC-AGI. She has joined Microsoft as a Principal researcher.

0 favorites 0 likes
#arc-agi

Accelerating Returns and the Qualitative Engine for Science

arXiv cs.AI · 2026-06-26 Cached

This paper examines Ray Kurzweil's thesis of accelerating returns and argues that while quantitative capabilities may accelerate, genuine scientific discovery requires a different capacity: qualitative reasoning about conceptual frameworks. It proposes the Qualitative Engine for Science (QES) as a response to this gap.

0 favorites 0 likes
#arc-agi

@rohanpaul_ai: GLM-5.2 got 22.8% on ARC-AGI-2:, $0.25/task To note here, around May 2025, the best verified models on ARC-AGI-2 were o…

X AI KOLs Timeline · 2026-06-24 Cached

GLM-5.2 achieves 22.8% on ARC-AGI-2 and 77% on ARC-AGI-1 at a low cost of $0.25 per task, representing a 7.6x improvement over the best frontier score from May 2025.

0 favorites 0 likes
#arc-agi

Sakana Fugu (3 minute read)

TLDR AI · 2026-06-22 Cached

Sakana AI introduces AB-MCTS, an inference-time scaling algorithm that enables multiple frontier AI models (Gemini 2.5 Pro, o4-mini, DeepSeek-R1-0528) to cooperate, significantly outperforming individual models on the ARC-AGI-2 benchmark.

0 favorites 0 likes
#arc-agi

Claude Opus 4.8 scores over 1% on ARC-AGI 3 !!

Reddit r/singularity · 2026-06-01

Claude Opus 4.8 achieves a score of over 1% on the ARC-AGI 3 benchmark, demonstrating slight progress on a difficult AI reasoning test.

0 favorites 0 likes
#arc-agi

@askalphaxiv: A fascinating paper supervised by Yoshua Bengio "Generative Recursive Reasoning" Test time compute should scale not jus…

X AI KOLs Timeline · 2026-05-21 Cached

The paper 'Generative Recursive Reasoning' introduces a method that scales test-time compute by sampling multiple latent reasoning trajectories in parallel, enabling the model to explore diverse hypotheses and avoid deterministic collapse. This approach improves performance on tasks such as Sudoku, ARC AGI, N Queens, and graph coloring, and can also generate valid Sudoku boards and MNIST digits.

0 favorites 0 likes
#arc-agi

Seed IQ ARC-AGI 3 Claims

Reddit r/ArtificialInteligence · 2026-05-15

A Reddit user debunks claims from Seed IQ (AGX) about solving the ARC-AGI-3 benchmark with a perfect score, arguing that refusal to submit to the Kaggle leaderboard (which allows closed-source submission) suggests a scam.

0 favorites 0 likes
#arc-agi

Useful Memories Become Faulty When Continuously Updated by LLMs

arXiv cs.AI · 2026-05-14 Cached

This paper shows that continuously consolidating past experiences into textual memory using LLMs degrades memory utility over time, and that preserving raw episodic trajectories outperforms forced consolidation, with implications for robust agentic memory systems.

0 favorites 0 likes
#arc-agi

Seed IQ-ARC AGI 3: Special behind-the-scenes look at Seed IQ on ARC-AGI 3 games! 14/14 games with a perfect 100% score across all.

Reddit r/ArtificialInteligence · 2026-05-13

Seed IQ achieves a perfect 14/14 score on ARC-AGI-3 games using an active inference, physics-driven multi-agent autonomous control engine, as shown in a behind-the-scenes video walkthrough.

0 favorites 0 likes
#arc-agi

Useful Memories Become Faulty When Continuously Updated by LLMs

Hugging Face Daily Papers · 2026-05-13 Cached

A study finds that continuously updating consolidated memories in LLM-based agentic systems degrades performance, and that retaining raw episodic trajectories is more reliable. Experiments on ARC-AGI show that even GPT-5.4 fails more often after consolidation.

0 favorites 0 likes
#arc-agi

@dylan_works_: Wrote up something fun I’ve been poking at: when LLM agents repeatedly rewrite their own experiences into textual “less…

X AI KOLs Timeline · 2026-05-09 Cached

This research blog post demonstrates that repeatedly rewriting LLM agent experiences into textual 'lessons' often degrades performance rather than improving it. The author finds that episodic memory retention performs better than abstract consolidation across various benchmarks like ARC-AGI and ALFWorld.

0 favorites 0 likes
#arc-agi

11.67% ARC-AGI-2 Local Eval on a Single 4090: The TOPAS Recursive Architecture

Reddit r/LocalLLaMA · 2026-05-07

The authors present TOPAS, a recursive AI architecture achieving 11.67% on ARC-AGI-2 using a single RTX 4090, aiming to demonstrate that architectural efficiency can outweigh raw compute power.

0 favorites 0 likes
← Previous
← Back to home

Submit Feedback