Tag
AMD upstreamed optimizations to PyTorch/TorchTitan and TorchAO for FP8 training on AMD Instinct GPUs, achieving up to 13.4% throughput gains on Llama3-8B and recovering 89% of FP8 quantization overhead on DeepSeek-V3 via fused Triton kernels.
The article describes the porting of PyTorch Monarch, a distributed training runtime, to AMD GPUs with ROCm, enabling single-controller fault-tolerant training at scale and addressing reliability challenges in large-scale LLM training.
Explores synthetic data generation, multi-agent optimization, and reinforcement learning to improve language models' ability to generate high-performance HIP kernels for AMD GPUs, demonstrating improvements in compilation and correctness rates on MI350X.
Zyphra releases ZAYA1-74B-Preview, a 74-billion parameter base model trained on AMD hardware, highlighting strong pre-RL reasoning capabilities and agentic performance signals.