Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?
Summary
This paper investigates whether coding agents require executable world models, simplification, and verification to solve the ARC-AGI-3 benchmark, contributing to research on AGI and reasoning.
View Cached Full Text
Cached at: 07/20/26, 09:21 AM
# Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3? Source: [https://arxiv.org/abs/2607.15439](https://arxiv.org/abs/2607.15439) Bibliographic Tools ## Bibliographic and Citation Tools Bibliographic Explorer Toggle Code, Data, Media ## Code, Data and Media Associated with this Article Demos ## Demos Related Papers ## Recommenders and Search Tools About arXivLabs ## arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website\. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy\. arXiv is committed to these values and only works with partners that adhere to them\. Have an idea for a project that will add value for arXiv's community?[**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html)\.
Similar Articles
ARCANA: A Reflective Multi-Agent Program Synthesis Framework for ARC-AGI-2 Reasoning
ARCANA is a reflective multi-agent framework that decomposes ARC-AGI-2 abstract reasoning tasks into iterative perception, hypothesis generation, symbolic execution, and reflective refinement, improving reasoning efficiency under strict constraints.
Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1
This paper presents cost-effective agent harnesses for ARC-AGI-1 that achieve strong performance using DeepSeek V3.2 without fine-tuning, via an Explorer-Definer Pipeline and a Reflective Orchestrator, achieving 67.25% pass@2 at low cost.
Coding Agent Is Good As World Simulator
This paper presents an agentic framework that uses coding agents to generate physically plausible world simulations from natural language prompts, outperforming video-based models in physical accuracy and instruction fidelity.
AGI Maze as a Benchmark Framework for World-Modeling Agents
This paper proposes the AGI Maze, a benchmark framework designed to evaluate the world-modeling capabilities of AI agents.
The Verification Horizon: No Silver Bullet for Coding Agent Rewards
This paper explores the challenges of verifying AI coding agents' outputs, arguing that verification is becoming harder than generation as models improve. It analyzes four reward constructions and shows that no fixed reward function remains effective as model capability grows.