@jino_rohit: after a lot of advice, im going to start to pick up blackwell kernels more seriously. ive mostly worked on server side …

X AI KOLs Timeline News

Summary

after a lot of advice, im going to start to pick up blackwell kernels more seriously. ive mostly worked on server side and model related optimizations and pre hopper kernels. now is a good time to learn some blackwell kernels!

after a lot of advice, im going to start to pick up blackwell kernels more seriously. ive mostly worked on server side and model related optimizations and pre hopper kernels. now is a good time to learn some blackwell kernels!
Original Article

Similar Articles

Blackwell and PDL performance increase

Reddit r/LocalLLaMA

Llama.cpp now supports Nvidia's Programmatic Dependent Launch (PDL) for Blackwell GPUs, offering a 5-10% performance boost on token generation. The feature is not enabled by default and requires a build flag.

@elliotarledge: For those wondering why I use a Kimi Linear megakernel instead of Qwen 3.6, first look at the parameter counts. One is …

X AI KOLs Timeline

Elliot Arledge explains why he prefers using a Kimi Linear megakernel over Qwen 3.6 for kernel performance, comparing parameter counts, layer synchronization, hidden dimensions, and architecture-specific optimizations. The discussion highlights that Kimi Linear architecture is more suitable for megakernel implementation, especially for batch-1 decode on RTX PRO 6000 Blackwell.