custom-kernels

Tag

Cards List
#custom-kernels

We don't need no stinkin' tensor library: solving poker in custom WebGPU kernels

Hacker News Top · 2026-07-30 Cached

The author details how they built an in-browser poker solver by using LLMs to generate custom WebGPU kernels instead of relying on a general tensor library, achieving over 10x speedup and demonstrating a paradigm shift where cheap generation can replace library abstraction.

0 favorites 0 likes
#custom-kernels

@no_stp_on_snek: busy evening in custom rust kernel and swift dispatch land. 3 of the 9 PRs are genuine new or changed metal kernels. th…

X AI KOLs Timeline · 2026-07-07 Cached

在自定义Rust内核和Swift调度逻辑的优化下,Qwen3.6-35b-A3B模型的预填速度在2k提示下从255 tok/s提升到1058 tok/s,实现了约4倍的加速,解码和困惑度未受影响。

0 favorites 0 likes
#custom-kernels

LFM2.5 230M running in-browser at 1,400 tok/s using custom WebGPU kernels

Reddit r/LocalLLaMA · 2026-06-25

LFM2.5 230M model achieves 1,400 tokens per second in-browser using custom WebGPU kernels, demonstrating efficient local inference.

0 favorites 0 likes
#custom-kernels

Running DeepSeek-V4 locally with 4x legacy RTX 2080 Ti ($2k budget setup). Custom Turing kernels, W8A8 quantization, and 255 prefill tok/s!

Reddit r/LocalLLaMA · 2026-05-20

A developer successfully runs DeepSeek-V4-Flash (284B total, 13B active) locally on four RTX 2080 Ti GPUs with a $2,500 budget, achieving 255 prefill tokens/s using custom Turing CUDA kernels, W8A8 quantization, and heterogeneous inference. The implementation is open-sourced.

0 favorites 0 likes
← Back to home

Submit Feedback