gpu-kernel-optimization

Tag

Cards List
#gpu-kernel-optimization

Optimizing CUDA like a Human: Micro-Profiling Tools as Expert Surrogates for LLM-Based GPU Kernel Optimization

arXiv cs.LG · 2026-06-26 Cached

KernelPro is a closed-loop multi-agent system that uses LLMs and micro-profiling tools to automatically optimize GPU kernel code, achieving geomean speedups of 2.42×/4.69×/5.30× on KernelBench and demonstrating a measured 11.6% energy reduction at matched speed.

0 favorites 0 likes
#gpu-kernel-optimization

First Steps Toward Automated AI Research (12 minute read)

TLDR AI · 2026-06-12 Cached

Recursive releases an automated AI research system that achieves state-of-the-art results on three benchmarks: fixed-budget language model training, small-model training speed, and GPU kernel optimization. The system automates the research loop and open-sources artifacts from its runs.

0 favorites 0 likes
#gpu-kernel-optimization

@Recursive_SI: https://x.com/Recursive_SI/status/2064980090702962699

X AI KOLs Timeline · 2026-06-11 Cached

Recursive releases early results from its automated AI research system, achieving state-of-the-art in fixed-budget language model training, small-model training speed, and GPU kernel optimization, and open-sources artifacts.

0 favorites 0 likes
#gpu-kernel-optimization

AgentKernelArena: Generalization-Aware Benchmarking of GPU Kernel Optimization Agents

Hugging Face Daily Papers · 2026-05-16 Cached

AgentKernelArena is an open-source benchmark for evaluating AI coding agents on GPU kernel optimization, assessing full agent workflows and generalization to unseen configurations across 196 tasks.

0 favorites 0 likes
← Back to home

Submit Feedback