@yminsky: We gave @dwarkesh_sp a tour of one of our new GPU-filled data-centers. Much fun!
Summary
Jane Street's tech lead Ron Minsky and physical engineering team lead Daniela Corvo gave a tour of a Texas data center retrofitted for high-density liquid-cooled GPU training, detailing liquid cooling systems, power distribution challenges, leak detection, and the opportunity cost of compute resources.
View Cached Full Text
Cached at: 05/16/26, 01:18 PM
We gave @dwarkesh_sp a tour of one of our new GPU-filled data-centers. Much fun!
https://t.co/WpmrVd46bf
TL;DR: Jane Street’s Head of Technology Ron Minsky and Head of Physical Engineering Daniela Corvo led a tour of a Texas data center retrofitted for high-density liquid-cooled GPU training, detailing liquid cooling systems, power distribution challenges, leak detection, and the opportunity cost of compute resources.
Data Center Overview: Transition from Air Cooling to Liquid Cooling
This Texas training data center was originally not designed for the current high-power-density racks. Traditional air-cooled cabinets typically range from 10-40 kW, while new GB300 cabinets peak at around 140 kW. The engineering team retrofitted the facility to introduce liquid cooling capability. Currently, about 85%-90% of the thermal load is removed via cold plates, with the remaining 15% handled by air cooling.
The key to liquid cooling is delivering 18°C coolant through quick-connects at the back to the cold plates on top of the GPUs, with the return flow at a higher temperature. Each sled automatically connects to the liquid supply, return, and 54V power when inserted. This design allows high-density racks to be deployed in the existing space—though total power capacity is limited by grid allocation, high-density layout saves space and even frees up areas in the hall for other uses (such as a podcast studio).
Engineering Details of the Liquid Cooling System
Leak Detection and Safety
Water has traditionally been prohibited in data centers, but liquid cooling introduces leakage risk. Leak detection ropes are placed under the floor inside racks, and if liquid is detected, the management switch triggers an alarm. Leak detection devices are also installed under the floor, with valves that can isolate the affected section. Leaks are uncommon, but as a new technology, long-term reliability remains to be seen.
Coolant and Heat Exchange
The building’s chilled water loop (around 18°C) exchanges heat via a heat exchanger with the internal “process water loop.” The process water must be distilled or deionized water mixed with 25% propylene glycol, filtered to 25 microns to prevent bacterial or algae growth that could clog cold plates. Propylene glycol inhibits microbial growth while ensuring efficient heat transfer.
Flow Balancing
Each rack has a valve and ultrasonic flow meter at the front, with a preset flow rate based on the maximum heat load. Precision adjustment ensures balanced coolant distribution from the beginning to the end of the pipe run.
Power System and Distribution Challenges
Circuit Breakers and Load Management
Power is distributed via busways, and the number of racks on each busway is strictly controlled to prevent overload trips. Breaker panels distribute power to different paths. The data center over-provisions the power distribution system to allow moving power between rows as CPU or GPU demands change. However, power flexibility is less than cooling: cooling can be adjusted by upsizing pipes, while power is limited by breaker current ratings, requiring precise load planning.
Peak Load Control
NVIDIA introduced a load management system in the new cabinets, using larger capacitors and software optimization to bring peak loads closer to average, keeping the curve flat. Jane Street’s own monitoring system reads data from breakers, understands the topology, and can automatically shut down power-hungry workloads if necessary to avoid tripping. Because hardware is extremely valuable, the operational strategy is to run as close to the limit as possible (over-subscription) while retaining safety mechanisms for a controlled fallback.
Cost of Interruption
If a breaker trips due to excessive current, running training tasks may be interrupted and need to recover from checkpoints, incurring enormous opportunity cost.
Opportunity Cost of Compute Resources
Ron Minsky emphasized that in an environment where compute resources are relatively inelastic, opportunity cost often dominates hardware cost. New compute resources take time to come online, and internal competition makes compute extremely expensive. Therefore, the data center pre-builds some headroom so that more GPUs can be quickly deployed when needed by the business.
Network Cabling: Fiber and Copper
High-density deployment brings cabling challenges. The entire facility uses about 8,000 km of fiber. Interestingly, most cables outside the cages are fiber, but the fastest internal connections use copper because electrons travel faster in copper than light in fiber (in terms of signal propagation speed). Latency optimization is applied across different network rates.
Scale of Cooling Infrastructure
The liquid cooling system includes buffer tanks (acting as thermal batteries) to cover the gap between a power outage and chiller restart, storing energy to keep cooling the GPUs while also buffering temperature fluctuations from workload changes. Traditional air-cooled units are still used in some areas, pulling hot air back, cooling it, and returning it to the hall. Wheel-operated valves are used to isolate leak points for repair—pulling a chain turns the valve to close or open.
From “Hive” to Modern Data Center
Twenty years ago, Jane Street’s first compute cluster, “the Hive,” was just six Dell boxes stacked at the end of a row. Today, the compute equipment themselves take up less space, while the infrastructure supporting them (transformers, chillers) has grown much larger. The data center has become a massive facility composed of both compute equipment and supporting infrastructure.
Source: https://www.youtube.com/watch?v=8J-GUnfSqeE
Similar Articles
@zostaff: 20 years ago Jane Street's entire compute cluster was six Dell boxes stacked on the floor at the end of an office row. …
Jane Street allowed Dwarkesh Patel to tour their new Texas data center with 4,032 GPUs, each rack pulling 140 kilowatts, highlighting the massive scale and unique networking choices.
@0xCheshire: Jane Street just released inside views of its Texas AI training center: 4,032 GPUs, 8,000 kilometers of fiber optic cable, and a fully deployed liquid cooling system because air cooling couldn't keep up. But what's truly stunning is the origin of this computing behemoth. Technical lead Ron Minsky recalls...
Jane Street revealed inside views of its AI training center in Texas, housing 4,032 GPUs, 8,000 kilometers of fiber optics, and a full liquid cooling system, while recounting the 20-year evolution from a humble start with six Dell hosts to today's extreme trading system.
@julien_c: May journalists everywhere read this
NVIDIA highlights a statistic from the Manhattan Institute showing that data centers account for only 0.2% of daily U.S. water usage, a figure that has declined in recent years due to new technologies.
@heyshrutimishra: Every data center needs cooling My Agent running on DGX Spark gets the full VIP treatment.
An AI agent running on NVIDIA's DGX Spark gets priority cooling treatment in a data center, highlighting infrastructure needs for AI workloads.
@apoorv03: One of the most substantive classes with @ChaseLochmiller at Stanford. We went deep on economics of the datacenter: - W…
Stanford class lecture by Chase Lochmiller dissecting the $650B AI infrastructure capex flow, margin capture, and shifting bottlenecks from GPUs to other datacenter constraints.