@realSharonZhou: We need better GPU performance benchmarks for RLing capabilities into frontier models for RSI, and I have a few ideas. …
Summary
Sharon Zhou proposes a vendor-agnostic, kernel-level GPU performance benchmark to help AI agents optimize compute efficiency for frontier models, and highlights AMD's AgentKernelArena as a starting point.
View Cached Full Text
Cached at: 07/05/26, 10:39 PM
We need better GPU performance benchmarks for RLing capabilities into frontier models for RSI, and I have a few ideas.
The goals would be to
-
help models improve their own compute efficiency (basically “living” RL envs with kernels that matter for new model generations AND new compute generations)
-
push performance with significantly less compute (useful in a compute crunch, also enables higher num iterations within a fixed compute budget; also more doable within existing RL setups)
-
invite more people to contribute to this space (by lowering the compute barrier to entry, including type of compute like not just GPUs/TPUs but even CPUs; but also different people like AI folk in addition to MLSys folk).
The idea is that we need a production-grade performance benchmark at the kernel level, that’s vendor-agnostic. It should measure the most useful kernels for top open weight LLMs, like GEMMs, attention, MoE, but also distributed ML stuff like collectives, so all-gather, dispatch/combine, etc. Make it interesting by including kernels that have strong differences between hardware vendors.
We have increasingly great benchmarks at the end-to-end model level (InferenceX @SemiAnalysis_ and Endpoints @MLPerf), but that takes significant compute to run useful things on, like online RL, or even to reproduce all configurations (precision, types of parallelism, etc) very frequently. Another benefit of kernel level perf, not full model perf, is we can also simulate a kernel on CPU, fairly quickly - so you could scale to different, cheaper compute surfaces.
A kernel level leaderboard can be updated in real time, much faster than the model performance stuff.
Here’s a start with AMD AgentKernelArena. I find that other AI researchers gravitate to this general setup. https://github.com/AMD-AGI/AgentKernelArena… A vendor-agnostic version would be fire and an incredibly useful asset to the community. It’d also really lower the compute barrier-to-entry to… chip in. Overall, I think this could pretty dramatically accelerate the field.
AMD-AGI/AgentKernelArena
Source: https://github.com/AMD-AGI/AgentKernelArena
AgentKernelArena: Competitive Arena for GPU Kernel Optimization Agents
As AI coding agents, like Claude Code and OpenAI Codex, rapidly improve, we need more than cherry-picked demos. Especially in specialized domains like GPU programming.
AgentKernelArena is a standardized evaluation arena built by AMD to measure how well AI coding agents perform on real GPU kernel optimization tasks.
Demo
A live illustrative demo is available at: http://165.245.130.75/
This demo is provided only to illustrate the results and should not be treated as the final benchmark leaderboard.
Overview & Features
AgentKernelArena provides an end-to-end, siloed benchmarking environment where LLM-powered agents (Cursor Agent, Claude Code, Codex, and custom agents) are evaluated side-by-side on the same kernel tasks using objective and reproducible metrics.
AgentKernelArena enables systematic evaluation of AI agents on GPU kernel optimization tasks by combining:
- Multi-Agent Arena: Cursor, Claude Code, Codex, and custom agents
- Multi-Model Support: OpenAI (GPT-5), Anthropic Claude (Opus and Sonnet families), and other models via OpenRouter or vLLM
- Task Categories: HIP (ROCm examples, rocPRIM, customer HIP), Triton (vLLM-style local harnesses and ROCmBench), and Torch2HIP conversions
- Real Metrics: Automated evaluation of compilation success, correctness, and real GPU performance speedups
- Designed for Fair Comparison: Standardized tasks, environments, prompts, and scoring for leaderboard-style evaluation
- A/B Testing for Agent Tools: Compare whether a new MCP server, skill, prompt, or agent-side tool actually improves outcomes by running the same task set with and without it and comparing standardized scores
- Workspace Isolation: Each task runs in a timestamped duplicate workspace for reproducibility
- Multi-GPU Parallel Runs: On multi-GPU servers, run one isolated Docker worker per GPU so idle GPUs immediately claim the next task from a shared queue
- Comprehensive Logging: Detailed logs with timestamps, prompts, outputs, and results for every task execution
- Flexible Configuration: YAML-based configuration for tasks, agents, and LLM parameters
A/B Testing and Ablation Studies
Beyond comparing different agents and models, AgentKernelArena can also be used to evaluate whether new agent-side capabilities actually help. For example, if you introduce a new MCP server, skill, prompt strategy, or tool integration, you can run the same task set twice — once with the capability enabled and once without it — and compare compilation, correctness, performance, and overall scores under the same evaluation conditions.
This makes AgentKernelArena useful not only as a leaderboard-style benchmark, but also as a controlled A/B testing framework for measuring the real impact of agent improvements.
Leaderboard Coming: Stay Tuned!
AgentKernelArena is actively under development. Upcoming releases will publish detailed evaluation results comparing agent performance across multiple task categories, using standardized correctness and performance scores.
| Model | Compiled | Correctness | Performance | Score |
|---|---|---|---|---|
| Cursor Agent | xx | xx | xx | xx |
| Claude Code | xx | xx | xx | xx |
| OpenAI Codex | xx | xx | xx | xx |
Architecture
Core Components
AgentKernelArena/
├── main.py # Main orchestration entry point
├── config.yaml # Global configuration
├── src/
│ ├── module_registration.py # Dynamic agent/prompt/post-processing loading
│ ├── preprocessing.py # Workspace setup and environment checks
│ ├── prompt_builder.py # Task prompt construction
│ ├── postprocessing.py # Result analysis and report generation
│ ├── score.py # Scoring logic for evaluation metrics
│ ├── tasks.py # Task discovery and registration
│ └── utils/
│ └── report_generation.py # Aggregate report analysis utilities
├── agents/
│ ├── cursor/ # Cursor agent integration
│ ├── claude_code/ # Claude Code agent integration
│ ├── codex/ # Codex CLI agent integration
│ ├── task_validator/ # Task quality validator
│ └── __init__.py # Agent registry
└── tasks/ # Task definitions
├── rocm-examples/ # ROCm example kernels
├── rocprim/ # rocPRIM kernels
├── customer_hip/ # Custom HIP kernels
├── triton/ # Triton benchmark kernels
└── torch2hip/ # Torch2HIP conversion tasks
Execution Flow
- Configuration Loading: Load
config.yamlwith agent, task, and LLM settings - Agent Registration: Dynamically load agent launcher, prompt builder, and post-processing handler based on AgentType enum
- Task Discovery: Scan
tasks/directory for task configurations matching specified categories - Workspace Setup: Create isolated workspace with timestamp for each task
- Prompt Building: Construct task-specific prompts from config, source code, and instructions/cheatsheets
- Agent Execution: Launch agent in workspace with constructed prompt
- Result Collection: Save agent output, logs, and modified code
- Post-Processing: Run compilation, correctness tests, performance profiling, and scoring
- Report Generation: Generate comprehensive evaluation report with metrics
For multi-GPU runs, the host-side Docker runner creates a shared .parallel/
queue under the run directory and starts one long-lived worker container per GPU.
Each worker claims tasks with an atomic descriptor move, runs one task at a time
with only one GPU visible, and then claims the next task. Final post-processing
runs once after all workers finish.
Installation
Prerequisites
- Docker
- The SGLang Docker image for your GPU arch (
gfx942useslmsysorg/sglang:v0.5.12-rocm720-mi30x;gfx950useslmsysorg/sglang:v0.5.12-rocm720-mi35x) - Git
- Host-installed agent CLIs for the agents you plan to evaluate
Setup
# Clone the repository
git clone <repository-url>
cd AgentKernelArena
# Install agent CLIs (examples)
# For Claude Code:
npm install -g @anthropic-ai/claude-code
# For Codex CLI: install per the official Codex CLI instructions,
# then ensure `codex` is available in PATH.
# Verify Docker can see the GPU/runtime and reuse agent login state.
make docker-smoke
make docker-check-agents
Performance timing helpers are maintained in src/tools/perf/; committed task
sources keep stubs that are materialized into run workspaces. See
src/tools/perf/README.md for the helper workflow.
Usage
Basic Usage
- Configure
config.yaml:
# Select agent type
agent:
template: claude_code # Options: cursor, claude_code, codex, task_validator
max_iterations: 5
# Specify tasks to run
tasks:
- rocm-examples/bitonic_sort
- customer_hip/silu
# - all # Run ALL tasks
target_gpu_model: MI300
log_directory: logs
workspace_directory_prefix: workspace
- Run evaluation:
make docker-run CONFIG=config.yaml
Run the same task set in parallel across multiple GPUs:
# Use explicit GPU IDs
make docker-parallel-run CONFIG=config.yaml GPU_IDS=0,1,2,3,4,5,6,7
# Or omit GPU_IDS to discover GPUs with rocm-smi --showid
make docker-parallel-run CONFIG=config.yaml
# Label the run directory
make docker-parallel-run CONFIG=config.yaml GPU_IDS=0,1 RUN_ARGS="--run-suffix parallel_smoke"
docker-parallel-run starts one Docker worker per GPU. Each worker gets isolated
agent state and cache directories, and sees a single logical GPU (0) inside the
container. It supports normal optimization agents and task_validator.
Resume a parallel run the same way as a serial run:
make docker-parallel-run CONFIG=config.yaml GPU_IDS=0,1,2,3 RUN_ARGS="--resume-latest"
make docker-parallel-run CONFIG=config.yaml GPU_IDS=0,1,2,3 RUN_ARGS="--resume-run run_20260702_041903_parallel8"
Advanced Usage
Running Specific Task Categories
tasks:
- rocm-examples/* # All ROCm examples
- rocprim/* # All rocPRIM tasks
- customer_hip/mmcv/* # All MMCV HIP kernels
- triton2triton/vllm/* # vLLM-style Triton kernel tasks
- triton2triton/rocmbench/* # ROCmBench Triton tasks
- instruction2triton/rocmbench/* # Instruction-to-Triton ROCmBench tasks
- torch2hip/* # All Torch2HIP conversion tasks
Task Configuration
Each task is defined by a config.yaml in its directory:
# tasks/rocm-examples/bitonic_sort/config.yaml
source_file_path:
- main.hip
target_kernel_functions:
- bitonic_sort_kernel
compile_command:
- make
correctness_command:
- ./applications_bitonic_sort -l 15
performance_command:
- rocprof-compute profile -n kernelgen --path rocprof_compute_profile --no-roof --join-type kernel -b SQ -b TCP -b TCC -- ./applications_bitonic_sort -l 15
- rocprof-compute analyze --path rocprof_compute_profile -b 2
task_type: hip2hip
prompt:
source_code: null # Optional: override default source code inclusion
instructions: null # Optional: custom instructions
cheatsheet: null # Optional: provide cheatsheet/reference
Scoring System
AgentKernelArena uses a cumulative scoring system:
| Metric | Points | Description |
|---|---|---|
| Compilation | 20 | Code compiles successfully without errors |
| Correctness | 100 | Code produces correct output (passes tests) |
| Speedup | ratio × 100 | Performance improvement over baseline |
Example: A submission that compiles (20), passes correctness (100), and achieves 1.5× speedup (150) would score 270 points.
Note: This is not the only way to score. Users could always define their own ways.
Development
Adding a New Agent
-
Create agent directory:
agents/your_agent/ -
Implement launch function:
# agents/your_agent/launch_agent.py
from agents import register_agent
@register_agent("your_agent")
def launch_agent(prompt: str, log_directory: str, workspace: str) -> str:
"""
Launch your agent.
Returns:
str: Agent output
"""
# Your agent implementation
return result
- Register in module_registration.py:
# Add to AgentType enum
class AgentType(Enum):
YOUR_AGENT = "your_agent"
# Add import in load_agent_launcher
if agent_type == AgentType.YOUR_AGENT:
from agents.your_agent import launch_agent
- Add prompt builder support (if needed):
# In load_prompt_builder
if agent_type in [..., AgentType.YOUR_AGENT]:
return prompt_builder
- Add post-processing support (if needed):
# In load_post_processing_handler
if agent_type in [..., AgentType.YOUR_AGENT]:
return general_post_processing
Adding a New Task
-
Create task directory:
tasks/<task_type>/<task_name>/ -
Add source files and scripts following this structure:
tasks/<task_type>/<task_name>/
├── config.yaml # Task configuration (required)
├── scripts/
│ └── task_runner.py # Compile/correctness/performance runner (recommended)
└── src/
└── <kernel files> # .cu, .hip, .py, etc.
- Create
config.yamlwith all required fields as lists (not scalar strings):
source_file_path:
- src/my_kernel.hip
target_kernel_functions:
- my_kernel_function
compile_command:
- python3 scripts/task_runner.py --mode compile
correctness_command:
- python3 scripts/task_runner.py --mode correctness
performance_command:
- python3 scripts/task_runner.py --mode performance
task_type: hip2hip # one of: hip2hip, cuda2hip, triton2triton, torch2hip, instruction2triton, repository, flydsl2flydsl
prompt:
source_code: null
instructions: null
cheatsheet: null
-
Add baseline performance (optional): Create
baseline.txtwith expected performance metrics -
Run the Task Validator Agent (required):
All new tasks must pass the task validator agent before being merged. The validator runs 10 automated checks covering config schema, source file existence, kernel symbol resolution, compilation, correctness, performance, self-containedness, GPU hang detection, correctness implementation review, and result template compatibility.
# Configure the validator to target your new task
# In config.yaml at repo root:
agent:
template: task_validator
tasks:
- <task_type>/<task_name>
# Run validation
make docker-run CONFIG=config.yaml
Review the generated validation_report.yaml in the workspace directory. The task must achieve PASS overall status (all checks pass). A WARN status (no failures but warnings) is acceptable with justification. A FAIL status means the task must be fixed before merging.
See agents/task_validator/README.md for the full list of validation checks and requirements.
Next Steps
- Enhance A/B Testing with Better Interactivity and User Experience
- Benchmarking State-of-the-Art Agents for Technical Reporting
- Standardize Holdout Tests with Comprehensive Shape Coverage
- Add Holdout Test Evaluation via Independent Agent
- New Feature: Support Multi Agents in Multi GPUs Server
- New Feature: Resume the Evaluation From Previous Experiment
- Agents Can Hang During Task Execution, Blocking Test Completion
- Expand Pytorch2HIP Task Set to 100+ Tasks
- Expand CUDA2HIP Task Set to 100+ Tasks
- Expand Triton2Triton Task Set to 100+ Tasks
- Expand HIP2HIP Task Set to 100+ Tasks
- Restructure Task Directory by Take Type and Difficulty Level
Similar Articles
AgentKernelArena: Generalization-Aware Benchmarking of GPU Kernel Optimization Agents
AgentKernelArena is an open-source benchmark for evaluating AI coding agents on GPU kernel optimization, assessing full agent workflows and generalization to unseen configurations across 196 tasks.
AA-AgentPerf-Local: Benchmarking local AI agents on laptops and workstations
Artificial Analysis open-sources AA-AgentPerf-Local, an inference benchmarking tool that replays real agent trajectories to measure how fast agentic AI runs on laptops and workstations, with initial results for NVIDIA DGX Spark, RTX 5090, AMD Ryzen AI Halo, and MacBook Pro M5 Pro.
@rohanpaul_ai: NVIDIA just posted the first agentic AI benchmark results where GB300 NVL72 runs up to 20x more coding agents per megaw…
NVIDIA published the first agentic AI benchmark results showing the GB300 NVL72 can run up to 20x more coding agents per megawatt than the H200, using the AgentPerf benchmark from Artificial Analysis.
(Rant ;)) Make your benchmarks realistic
A community rant urging realistic AI model benchmarks that account for context size, multimodal features, hardware specifics, and parallel processing, rather than just raw speed.
@TheAhmadOsman: Local AI is now good btw
Ahmad announces a Local AI Hardware Arena using ODS to benchmark LLMs on hardware like RTX PRO 6000, DGX Spark, Strix Halo, M5 MacBook Pro, and ChatGPT, inviting community input for future comparisons.