Tag
This paper explores approximating softmax in pretrained LLMs for kernel acceleration, demonstrating performance gains like up to 25.8% speedup on Blackwell B200 with minimal perplexity impact.
AMD upstreamed optimizations to PyTorch/TorchTitan and TorchAO for FP8 training on AMD Instinct GPUs, achieving up to 13.4% throughput gains on Llama3-8B and recovering 89% of FP8 quantization overhead on DeepSeek-V3 via fused Triton kernels.
SKT released A.X K2, a 688B-parameter sparse MoE language model with 33B active parameters, natively trained in FP8 and featuring Think/Non-Think reasoning modes, on Hugging Face.