kernel-optimization

Tag

Cards List
#kernel-optimization

@runinfrai: we spent september rewriting the kernels behind GLM 5.3 Flash. major release is live on RunInfra today 670 tok/s on Ver…

X AI KOLs Timeline ↗ · 14h ago Cached

RunInfra announces a major release of GLM 5.3 Flash with rewritten inference kernels, delivering 670 tok/s on an FP8, vendor-native build now also running on AMD, with competitive pricing and a 1M token context window.

0 favorites 0 likes
#kernel-optimization

Splash on a 40-core M5 Max: +20% decode by tuning the kernels for your own chip

Reddit r/LocalLLaMA ↗ · 5d ago

The article describes how to use the 'make tune-kernels' tool in Splash to optimize AI model decoding on 40-core M5 Max chips, achieving up to 20% speed improvement by tuning kernel layouts for specific hardware.

0 favorites 0 likes
#kernel-optimization

Show HN: Agentic CUDA Kernel Optimizer

Hacker News Top ↗ · 5d ago Cached

An agentic CUDA kernel optimizer that automates GPU implementation generation through iterative code generation, correctness checks, benchmarking, and refinement, powered by LangGraph and OpenAI models.

0 favorites 0 likes
#kernel-optimization

@RisingSayak: Found a faster kernel? You shouldn’t need to rewrite your model to use it. With Kernels, you can choose which kernel ru…

X AI KOLs Following ↗ · 2026-09-18 Cached

This Twitter thread introduces Hugging Face's Kernels, a tool that allows users to select and replace optimized kernel implementations for supported layers in AI models without rewriting the entire model.

0 favorites 0 likes
#kernel-optimization

@no_stp_on_snek: DeepSeek-V4.1-Flash on 2 Sparks. thinking high, TP=2. 55M tokens. Actual kernel work, not a demo. it read the notes and…

X AI KOLs Timeline ↗ · 2026-09-15 Cached

The article describes an optimization for the DeepSeek-V4.1-Flash AI model on Metal hardware, where the attention kernel was improved to only launch necessary tiles, reducing compute waste and enhancing inference speed.

0 favorites 0 likes
#kernel-optimization

@reprompting: reading about tritonBLAS today https://arxiv.org/pdf/2512.04226

X AI KOLs Timeline ↗ · 2026-09-14 Cached

The paper presents tritonBLAS, an analytical model for optimizing GPU GEMM kernel parameters without runtime autotuning, achieving near-optimal performance and significantly reducing compilation time.

0 favorites 0 likes
#kernel-optimization

AMDKernelVault: Large-Scale Datasets and Agentic Training for AMD GPU Kernel Optimization

arXiv cs.CL ↗ · 2026-09-14 Cached

The article introduces AMDKernelVault, an open dataset and training framework for AMD GPU kernel optimization, featuring large-scale HIP and Triton kernels and agent-driven pipelines for generating and validating kernels.

0 favorites 0 likes
#kernel-optimization

Google Accelerator Agents for TPU Development (GitHub Repo)

TLDR AI ↗ · 2026-09-08 Cached

Google Accelerator Agents is a GitHub repository of AI-powered tools to accelerate machine learning development on TPUs, featuring agents for code migration and kernel optimization using Gemini.

0 favorites 0 likes
#kernel-optimization

DataKernelBench: Can LLMs Optimize Database Queries on GPUs?

arXiv cs.CL ↗ · 2026-08-27 Cached

This paper introduces DataKernelBench, a benchmark for evaluating LLMs on optimizing GPU kernels for database queries, achieving speedups over baseline methods like torch.compile.

0 favorites 0 likes
#kernel-optimization

KernelArc: A Multi-Agent Framework for GPU Kernel Optimization

arXiv cs.AI ↗ · 2026-08-19 Cached

KernelArc is a multi-agent framework that uses strategy-specialized agents to autonomously optimize GPU kernels across heterogeneous workloads, achieving top rankings on NVIDIA GPU benchmarks.

0 favorites 0 likes
#kernel-optimization

CAKE: Compiler-Agent Co-Design for Frontier Kernel Evolution

arXiv cs.LG ↗ · 2026-08-14 Cached

This paper presents CAKE, a compiler-agent co-design framework that lets AI agents author a hardware-explicit IR for GPU kernels, achieving significant speedups over baselines across various workload families.

0 favorites 0 likes
#kernel-optimization

Kernel Forge: An Agent Harness for LLM-based Generation and Optimization of CUDA Kernels

arXiv cs.AI ↗ · 2026-07-29 Cached

Kernel Forge is an open-source agent harness that uses LLMs and Monte Carlo Tree Search to automatically generate and optimize CUDA kernels for any unmodified PyTorch model, achieving up to 2.83× speedup on softmax in Gemma 4 E2B.

0 favorites 0 likes
#kernel-optimization

Spent two weeks on a kernel that benchmarked 29x faster. End to end it's maybe 6-10%, and it's not even wired in yet.

Reddit r/LocalLLaMA ↗ · 2026-07-24

The author optimized a matmul kernel for BitNet's ternary models on CPU, achieving 29x speedup in isolation, but found that the model is memory-bound, resulting in only 6-10% end-to-end gain. The inference engine is available as open-source.

0 favorites 0 likes
#kernel-optimization

SonicSampler: Unified Tile-Aware Kernels for LLM Sampling and Speculative Verification

arXiv cs.AI ↗ · 2026-07-24 Cached

SonicSampler presents a unified suite of tile-aware Triton kernels that vertically fuse the entire LLM sampling pipeline, supporting dynamic per-request behaviors and speculative verification, achieving up to 16x speedup over state-of-the-art baselines.

0 favorites 0 likes
#kernel-optimization

JAXBench: Benchmarking Autonomous TPU Kernel Optimization

arXiv cs.AI ↗ · 2026-07-24 Cached

JAXBench is a new benchmark suite of 50 JAX workloads for evaluating AI-generated kernel optimization on Google Cloud TPUs, with hand-tuned baselines and an agent evaluation harness. The paper finds that conditioning on curated TPU documentation significantly improves correctness and speedup, with Autocomp beam-search achieving up to 1.6x geomean speedup over XLA on hand-tuned kernels.

0 favorites 0 likes
#kernel-optimization

@gpuwaster: 61/100 of GPU Grind going through this GTC 2020 lecture: Developing CUDA kernels to push Tensor Cores to the Absolute L…

X AI KOLs Timeline ↗ · 2026-07-11 Cached

The author, as part of a 100-part GPU learning series, reviews a GTC 2020 lecture by Andrew Kerr on developing high-performance CUDA kernels for Tensor Cores on NVIDIA A100, discussing techniques and the trade-off between raw CUDA and CUTLASS.

0 favorites 0 likes
#kernel-optimization

@realSharonZhou: We need better GPU performance benchmarks for RLing capabilities into frontier models for RSI, and I have a few ideas. …

X AI KOLs Timeline ↗ · 2026-07-05 Cached

Sharon Zhou proposes a vendor-agnostic, kernel-level GPU performance benchmark to help AI agents optimize compute efficiency for frontier models, and highlights AMD's AgentKernelArena as a starting point.

0 favorites 0 likes
#kernel-optimization

@reach_vb: I’ve got a ton of personal side projects I want to tackle with Sol: 1. Optimise MTP support in llama.cpp (to find low h…

X AI KOLs Timeline ↗ · 2026-07-02 Cached

The user shares plans to use Sol for optimizing MTP support in llama.cpp, standardizing chat templates, reviving an abandoned whisper project, finishing a personal finance app, and kernel optimization.

0 favorites 0 likes
#kernel-optimization

@googlegemma: “Agentic kernel optimization is the future of on-device inference” @xenovacom used Fable 5 to write kernels that pushed…

X AI KOLs Timeline ↗ · 2026-07-01 Cached

Xenova used Fable 5 to write optimized kernels achieving 255 tokens per second for Gemma 4 on WebGPU with M4, demonstrating agentic kernel optimization for on-device inference.

0 favorites 0 likes
#kernel-optimization

@h100envy: PyTorch core engineer at Meta turned CUDA kernel writing into a sport in 13 minutes - better than $1500 GPU programming…

X AI KOLs Timeline ↗ · 2026-06-30 Cached

A PyTorch core engineer at Meta demonstrated a fast CUDA kernel optimization loop that outperforms expensive bootcamps, with the winning code merged into PyTorch via the KernelBot competition.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback