@jino_rohit: after a lot of advice, im going to start to pick up blackwell kernels more seriously. ive mostly worked on server side …
Summary
after a lot of advice, im going to start to pick up blackwell kernels more seriously. ive mostly worked on server side and model related optimizations and pre hopper kernels. now is a good time to learn some blackwell kernels!
Similar Articles
@RisingSayak: The kernels project at Hugging Face has been growing! We want it to be the go-to place for kernel devs and kernel users…
Hugging Face's kernels project is expanding and seeking contributors for agentic kernel development to provide real optimization value to models.
Blackwell and PDL performance increase
Llama.cpp now supports Nvidia's Programmatic Dependent Launch (PDL) for Blackwell GPUs, offering a 5-10% performance boost on token generation. The feature is not enabled by default and requires a build flag.
@coffeecup2020: If your card support Blackwell, read this! https://github.com/turbo-tan/llama.cpp-tq3… updated with turbo4/turbo3 TQ3_4…
A llama.cpp fork introduces TurboQuant TQ3_4S quantization that maps to Blackwell FP4 tensor cores, achieving up to 221% faster prompt processing on GB10 while maintaining near Q4 quality at Q3 size.
@RisingSayak: Found a faster kernel? You shouldn’t need to rewrite your model to use it. With Kernels, you can choose which kernel ru…
This Twitter thread introduces Hugging Face's Kernels, a tool that allows users to select and replace optimized kernel implementations for supported layers in AI models without rewriting the entire model.
@elliotarledge: For those wondering why I use a Kimi Linear megakernel instead of Qwen 3.6, first look at the parameter counts. One is …
Elliot Arledge explains why he prefers using a Kimi Linear megakernel over Qwen 3.6 for kernel performance, comparing parameter counts, layer synchronization, hidden dimensions, and architecture-specific optimizations. The discussion highlights that Kimi Linear architecture is more suitable for megakernel implementation, especially for batch-1 decode on RTX PRO 6000 Blackwell.