@svpino: This is a cool benchmark that uses video games to test agents. Basically, they speedrun the game to test how well agent…

X AI KOLs Timeline Tools

Summary

SpeedrunBench is a new benchmark that evaluates AI agents by having them speedrun video games to test planning, learning, and decision-making, with a public leaderboard for model comparison.

This is a cool benchmark that uses video games to test agents. Basically, they speedrun the game to test how well agents can plan, act, learn from previous attempts, and optimize a long sequence of decisions. The goal is to evaluate agents' capability through an autoresearch-like loop There's a public leaderboard where you can see how good your favorite model is.
Original Article
View Cached Full Text

Cached at: 09/04/26, 02:19 AM

This is a cool benchmark that uses video games to test agents.

Basically, they speedrun the game to test how well agents can plan, act, learn from previous attempts, and optimize a long sequence of decisions.

The goal is to evaluate agents’ capability through an autoresearch-like loop

There’s a public leaderboard where you can see how good your favorite model is.

Anand Kannappan (@anandnk24): Today, we’re releasing SpeedrunBench: the first benchmark that measures how fast agents can beat video games.

We asked frontier models not to simply complete video games, but to speedrun them - across 10 titles like Mario Kart and Pokemon.

On simpler games, agents get close to

Similar Articles

ProgramBench (5 minute read)

TLDR AI

ProgramBench is a new benchmark that evaluates AI agents' ability to reconstruct complete software projects from compiled binaries and documentation without access to source code or decompilation tools.