flashinfer

Tag

Cards List
#flashinfer

@h100envy: CMU PhD who built the kernels NVIDIA now ships in TensorRT-LLM explained fast attention in 68 minutes - better than $12…

X AI KOLs Timeline · 2026-07-02 Cached

A CMU PhD who developed the kernels now used by NVIDIA in TensorRT-LLM explains fast attention, covering fused CUDA kernels, FlashInfer, Triton, and paged-KV attention, enabling more tokens per second on the same GPU.

0 favorites 0 likes
#flashinfer

@KeisukeKamahori: Very excited to share that our team at @UWSyFi won multiple prizes at the FlashInfer AI Kernel Generation Contest in #M…

X AI KOLs Following · 2026-05-22 Cached

University of Washington SyFI team won multiple prizes at the FlashInfer AI Kernel Generation Contest held during MLSys2026, with support from NVIDIA and Modal.

0 favorites 0 likes
← Back to home

Submit Feedback