Tag
RunInfra announces a major release of GLM 5.3 Flash with rewritten inference kernels, delivering 670 tok/s on an FP8, vendor-native build now also running on AMD, with competitive pricing and a 1M token context window.
The article describes how to use the 'make tune-kernels' tool in Splash to optimize AI model decoding on 40-core M5 Max chips, achieving up to 20% speed improvement by tuning kernel layouts for specific hardware.
An agentic CUDA kernel optimizer that automates GPU implementation generation through iterative code generation, correctness checks, benchmarking, and refinement, powered by LangGraph and OpenAI models.
This Twitter thread introduces Hugging Face's Kernels, a tool that allows users to select and replace optimized kernel implementations for supported layers in AI models without rewriting the entire model.
The article describes an optimization for the DeepSeek-V4.1-Flash AI model on Metal hardware, where the attention kernel was improved to only launch necessary tiles, reducing compute waste and enhancing inference speed.
The paper presents tritonBLAS, an analytical model for optimizing GPU GEMM kernel parameters without runtime autotuning, achieving near-optimal performance and significantly reducing compilation time.
The article introduces AMDKernelVault, an open dataset and training framework for AMD GPU kernel optimization, featuring large-scale HIP and Triton kernels and agent-driven pipelines for generating and validating kernels.
Google Accelerator Agents is a GitHub repository of AI-powered tools to accelerate machine learning development on TPUs, featuring agents for code migration and kernel optimization using Gemini.
This paper introduces DataKernelBench, a benchmark for evaluating LLMs on optimizing GPU kernels for database queries, achieving speedups over baseline methods like torch.compile.
KernelArc is a multi-agent framework that uses strategy-specialized agents to autonomously optimize GPU kernels across heterogeneous workloads, achieving top rankings on NVIDIA GPU benchmarks.
This paper presents CAKE, a compiler-agent co-design framework that lets AI agents author a hardware-explicit IR for GPU kernels, achieving significant speedups over baselines across various workload families.
Kernel Forge is an open-source agent harness that uses LLMs and Monte Carlo Tree Search to automatically generate and optimize CUDA kernels for any unmodified PyTorch model, achieving up to 2.83× speedup on softmax in Gemma 4 E2B.
The author optimized a matmul kernel for BitNet's ternary models on CPU, achieving 29x speedup in isolation, but found that the model is memory-bound, resulting in only 6-10% end-to-end gain. The inference engine is available as open-source.
SonicSampler presents a unified suite of tile-aware Triton kernels that vertically fuse the entire LLM sampling pipeline, supporting dynamic per-request behaviors and speculative verification, achieving up to 16x speedup over state-of-the-art baselines.
JAXBench is a new benchmark suite of 50 JAX workloads for evaluating AI-generated kernel optimization on Google Cloud TPUs, with hand-tuned baselines and an agent evaluation harness. The paper finds that conditioning on curated TPU documentation significantly improves correctness and speedup, with Autocomp beam-search achieving up to 1.6x geomean speedup over XLA on hand-tuned kernels.
The author, as part of a 100-part GPU learning series, reviews a GTC 2020 lecture by Andrew Kerr on developing high-performance CUDA kernels for Tensor Cores on NVIDIA A100, discussing techniques and the trade-off between raw CUDA and CUTLASS.
Sharon Zhou proposes a vendor-agnostic, kernel-level GPU performance benchmark to help AI agents optimize compute efficiency for frontier models, and highlights AMD's AgentKernelArena as a starting point.
The user shares plans to use Sol for optimizing MTP support in llama.cpp, standardizing chat templates, reviving an abandoned whisper project, finishing a personal finance app, and kernel optimization.
Xenova used Fable 5 to write optimized kernels achieving 255 tokens per second for Gemma 4 on WebGPU with M4, demonstrating agentic kernel optimization for on-device inference.
A PyTorch core engineer at Meta demonstrated a fast CUDA kernel optimization loop that outperforms expensive bootcamps, with the winning code merged into PyTorch via the KernelBot competition.