gpu-clusters

Tag

Cards List
#gpu-clusters

@nrehiew_: This was great! A fun followup is given some arbitrary topology and architecture, compute what is the expected compute/…

X AI KOLs Timeline ↗ · 2026-07-15 Cached

A tweet reply to a blog post about collective communication in TPU/GPU clusters, suggesting a follow-up to compute expected compute/communication duration for a 1T MoE on Blackwell NVL72 and estimate theoretical MFU and bottlenecks.

0 favorites 0 likes
#gpu-clusters

Nobody knows what a used GPU cluster is worth

Hacker News Top ↗ · 2026-07-15 Cached

An analysis of the challenges in valuing used GPU clusters as collateral for debt financing, highlighting operational dependencies and hardware failure rates that make traditional asset appraisal impossible.

0 favorites 0 likes
#gpu-clusters

GPU cluster sitting idle waiting on storage, more common than I expected

Reddit r/ArtificialInteligence ↗ · 2026-07-13

At an AI infrastructure meetup, multiple people reported that GPU clusters often sit idle because storage systems can't feed data fast enough during training, especially with large unstructured datasets on older NAS. Attendees mentioned moving to high-throughput S3-optimized platforms like Cloudian HyperStore and VAST Data to address the bottleneck.

0 favorites 0 likes
#gpu-clusters

@TheNoise2Signal: How does frontier training use 2,048 GPUs? Because there are five dimensions you can split work across - and at scale, …

X AI KOLs Timeline ↗ · 2026-05-25 Cached

Explains how frontier AI training uses up to 2,048 GPUs by splitting work across five dimensions, demystifying model training frameworks.

0 favorites 0 likes
#gpu-clusters

Unlocking large scale AI training networks with MRC (Multipath Reliable Connection)

OpenAI Blog ↗ · 2026-05-05 Cached

OpenAI has released MRC (Multipath Reliable Connection), a novel networking protocol developed with industry partners to improve performance and resilience in large-scale AI training clusters. The specification was published via the Open Compute Project to standardize infrastructure for efficient supercomputer operations.

0 favorites 0 likes
#gpu-clusters

Blowing Off Steam: How Power-Flexible AI Factories Can Stabilize the Global Energy Grid

NVIDIA Blog ↗ · 2026-03-25 Cached

Emerald AI demonstrated how power-flexible AI factories can autonomously adjust electricity consumption to stabilize grid demand, using NVIDIA GPUs and infrastructure at a London data center to absorb peak power surges without disrupting critical workloads.

0 favorites 0 likes
#gpu-clusters

Scaling Kubernetes to 7,500 nodes

OpenAI Blog ↗ · 2021-01-25 Cached

OpenAI shares detailed lessons learned from scaling a single Kubernetes cluster to 7,500 nodes to support large machine learning workloads, covering networking, scheduling, and infrastructure challenges. The post builds on their earlier experience scaling to 2,500 nodes and aims to help the broader Kubernetes community.

0 favorites 0 likes
← Back to home

Submit Feedback