Tag
The author built FlashMLA for consumer-grade Blackwell sm_120, achieving 2-3x performance gains over PyTorch SDPA in attention-heavy workloads like long-context training and sparse prefill.