kernel-acceleration

Tag

Cards List
#kernel-acceleration

Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration

Hugging Face Daily Papers ↗ · 4d ago Cached

This paper explores approximating softmax in pretrained LLMs for kernel acceleration, demonstrating performance gains like up to 25.8% speedup on Blackwell B200 with minimal perplexity impact.

0 favorites 0 likes
← Back to home

Submit Feedback