Tag
This article describes a method to train LoRA adapters using AsyncGRPOTrainer and sync them via Storage Buckets across separate Hugging Face Jobs, eliminating the need for NCCL communication.
NVIDIA shares debugging lessons from its Exemplar Cloud program, detailing how configuration issues in SMMU power management, NUMA placement, NCCL queue-pair concurrency, and hardware defects cause 8-12% training throughput gaps on AI clusters, and how to diagnose and fix them.