AGI Maze as a Benchmark Framework for World-Modeling Agents
Summary
This paper proposes the AGI Maze, a benchmark framework designed to evaluate the world-modeling capabilities of AI agents.
View Cached Full Text
Cached at: 07/02/26, 05:40 AM
# AGI Maze as a Benchmark Framework for World-Modeling Agents Source: [https://arxiv.org/abs/2607.00627](https://arxiv.org/abs/2607.00627) Bibliographic Tools ## Bibliographic and Citation Tools Bibliographic Explorer Toggle Code, Data, Media ## Code, Data and Media Associated with this Article Demos ## Demos Related Papers ## Recommenders and Search Tools About arXivLabs ## arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website\. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy\. arXiv is committed to these values and only works with partners that adhere to them\. Have an idea for a project that will add value for arXiv's community?[**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html)\.
Similar Articles
AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research
Introduces AutoWorldModel-Bench, a closed-loop benchmark for evaluating AI coding agents on autonomous world-model research across eight game environments. The benchmark shows frontier agents like Codex-5.4 and Claude Opus 4.6 make non-trivial research-style improvements in most sessions.
Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments
GauntletBench is a new web-based benchmark that evaluates AI agents on challenging scenarios focusing on temporal perception, graphical understanding, and 3D reasoning. Results show state-of-the-art agents achieve only 19.1% success rate compared to over 80% for non-expert humans, highlighting significant limitations in current agentic systems.
Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?
This paper investigates whether coding agents require executable world models, simplification, and verification to solve the ARC-AGI-3 benchmark, contributing to research on AGI and reasoning.
Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence
Agent-World introduces a self-evolving training framework for general agent intelligence that autonomously discovers real-world environments and tasks via the Model Context Protocol, enabling continuous learning. Agent-World-8B and 14B models outperform strong proprietary models across 23 challenging agent benchmarks.
Building World Models with Agent Swarms
The author describes building a world model harness that coordinates research agents to maximize evaluation metrics, achieving a 25x score improvement in a multimodal masked reconstruction task using geospatial inputs.