标签
AMD upstreamed optimizations to PyTorch/TorchTitan and TorchAO for FP8 training on AMD Instinct GPUs, achieving up to 13.4% throughput gains on Llama3-8B and recovering 89% of FP8 quantization overhead on DeepSeek-V3 via fused Triton kernels.