fp8-training

Tag

Cards List
#fp8-training

Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration

Hugging Face Daily Papers ↗ · 4d ago Cached

This paper explores approximating softmax in pretrained LLMs for kernel acceleration, demonstrating performance gains like up to 25.8% speedup on Blackwell B200 with minimal perplexity impact.

0 favorites 0 likes
#fp8-training

@PyTorch: AMD has been upstreaming optimizations for improved FP8 training support in PyTorch/TorchTitan and PyTorch/TorchAO, mak…

X AI KOLs Timeline ↗ · 2026-08-13 Cached

AMD upstreamed optimizations to PyTorch/TorchTitan and TorchAO for FP8 training on AMD Instinct GPUs, achieving up to 13.4% throughput gains on Llama3-8B and recovering 89% of FP8 quantization overhead on DeepSeek-V3 via fused Triton kernels.

0 favorites 0 likes
#fp8-training

@_akhaliq: A.X K2 just dropped on Hugging Face Large-Scale Sparse MoE (688B / 33B Active) https://huggingface.co/skt/A.X-K2

X AI KOLs Following ↗ · 2026-07-29 Cached

SKT released A.X K2, a 688B-parameter sparse MoE language model with 33B active parameters, natively trained in FP8 and featuring Think/Non-Think reasoning modes, on Hugging Face.

0 favorites 0 likes
← Back to home

Submit Feedback