VSArena — an open benchmark for AI agents in 3D embodied environments
Summary
VSArena is an open benchmark for evaluating AI agents in interactive 3D environments, offering remote execution and a public leaderboard to assess perception, reasoning, and action without physical robots.
Similar Articles
VSArena v0.6.0 — a new Studio for running and inspecting embodied AI policies in the browser
VSArena v0.6.0 is a major update to a browser-based studio for running and evaluating Vision-Language-Action policies in 3D physics environments, emphasizing reliable and reproducible assessment without local simulators or physical robots.
An AI agent just stacked blocks in a live physics simulation — building an open benchmark arena for embodied AI
An early-stage project is building an open, browser-based arena for AI agents to compete in real-time physical reasoning tasks, demonstrated by an AI agent achieving 100% task completion and high spatial accuracy in a block-stacking simulation.
Agent Arena
Agent Arena is the first public arena for AI agents, allowing users to test and compare AI agents in a competitive environment.
Inside VAKRA: Reasoning, Tool Use, and Failure Modes of Agents
This article introduces VAKRA, an executable benchmark for evaluating AI agents' reasoning and tool-use capabilities in enterprise-like environments. It analyzes failure modes and details the benchmark's structure involving API chaining and document retrieval.
VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control
VABench introduces a benchmark to evaluate embodied spatial intelligence in models by testing their ability to observe, reason, and act through visual demonstrations and active perception. It shows that active camera control improves task success, but no model completes long-horizon episodes.