Tag
A CMU PhD who developed the kernels now used by NVIDIA in TensorRT-LLM explains fast attention, covering fused CUDA kernels, FlashInfer, Triton, and paged-KV attention, enabling more tokens per second on the same GPU.
University of Washington SyFI team won multiple prizes at the FlashInfer AI Kernel Generation Contest held during MLSys2026, with support from NVIDIA and Modal.