@svpino: This is a cool benchmark that uses video games to test agents. Basically, they speedrun the game to test how well agent…
Summary
SpeedrunBench is a new benchmark that evaluates AI agents by having them speedrun video games to test planning, learning, and decision-making, with a public leaderboard for model comparison.
View Cached Full Text
Cached at: 09/04/26, 02:19 AM
This is a cool benchmark that uses video games to test agents.
Basically, they speedrun the game to test how well agents can plan, act, learn from previous attempts, and optimize a long sequence of decisions.
The goal is to evaluate agents’ capability through an autoresearch-like loop
There’s a public leaderboard where you can see how good your favorite model is.
Anand Kannappan (@anandnk24): Today, we’re releasing SpeedrunBench: the first benchmark that measures how fast agents can beat video games.
We asked frontier models not to simply complete video games, but to speedrun them - across 10 titles like Mario Kart and Pokemon.
On simpler games, agents get close to
Similar Articles
@seclink: 又一个 benchmark , 收录一下 样本和标注。
PatronusAI releases SpeedrunBench, the first benchmark that measures how fast AI agents can speedrun video games across multiple titles.
@ms_aifrontiers: SentinelBench tests agents in time-evolving web environments where success requires waiting. How you wait matters: on 4…
SentinelBench is a new benchmark for testing AI agents in time-evolving web environments. It finds that agents using a specialized change-detection tool outperform those using sleep-and-poll loops, reducing cost by 9.7x.
ProgramBench (5 minute read)
ProgramBench is a new benchmark that evaluates AI agents' ability to reconstruct complete software projects from compiled binaries and documentation without access to source code or decompilation tools.
Are we missing a benchmark for agent runtimes, not just models?
The article discusses the need for a benchmark to evaluate AI agent runtimes independently of models, suggesting metrics like task success rate and cost, and proposing controlled experiments to compare platforms.
@gneubig: SkillsBench is a great benchmark, and I'm not just saying that because @OpenHandsDev beats everyone else on it Skills a…
SkillsBench 1.1 is a benchmark for evaluating AI agents' ability to use skills, now fully audited and error-free.