Tag
This paper explores approximating softmax in pretrained LLMs for kernel acceleration, demonstrating performance gains like up to 25.8% speedup on Blackwell B200 with minimal perplexity impact.