Tag
A tweet reply to a blog post about collective communication in TPU/GPU clusters, suggesting a follow-up to compute expected compute/communication duration for a 1T MoE on Blackwell NVL72 and estimate theoretical MFU and bottlenecks.
An analysis of the challenges in valuing used GPU clusters as collateral for debt financing, highlighting operational dependencies and hardware failure rates that make traditional asset appraisal impossible.
At an AI infrastructure meetup, multiple people reported that GPU clusters often sit idle because storage systems can't feed data fast enough during training, especially with large unstructured datasets on older NAS. Attendees mentioned moving to high-throughput S3-optimized platforms like Cloudian HyperStore and VAST Data to address the bottleneck.
Explains how frontier AI training uses up to 2,048 GPUs by splitting work across five dimensions, demystifying model training frameworks.
OpenAI has released MRC (Multipath Reliable Connection), a novel networking protocol developed with industry partners to improve performance and resilience in large-scale AI training clusters. The specification was published via the Open Compute Project to standardize infrastructure for efficient supercomputer operations.
Emerald AI demonstrated how power-flexible AI factories can autonomously adjust electricity consumption to stabilize grid demand, using NVIDIA GPUs and infrastructure at a London data center to absorb peak power surges without disrupting critical workloads.
OpenAI shares detailed lessons learned from scaling a single Kubernetes cluster to 7,500 nodes to support large machine learning workloads, covering networking, scheduling, and infrastructure challenges. The post builds on their earlier experience scaling to 2,500 nodes and aims to help the broader Kubernetes community.