Data Centers

Liquid Cooling at 45°C: The New Thermal Math of AI Factories

Liquid Cooling at 45°C: The New Thermal Math of AI Factories

AI data center cooling has moved from a facilities concern to a hard limit on usable GPU capacity. NVIDIA says its Rubin AI infrastructure can accept cooling liquid at up to 45°C (113°F), with coolant leaving the system at roughly 55°C. That temperature window changes the thermal math: a site can reject more heat with dry coolers and less compressor work, while keeping dense accelerators inside their validated operating range.

The headline needs a caveat. A 45°C supply loop does not make every AI facility chiller-free. Outdoor design temperature, humidity, loop approach temperature, water chemistry, redundancy, and the server vendor warranty still decide the final plant. But it gives infrastructure teams a materially larger operating envelope than conventional chilled-water assumptions.

Why rack density broke the old model

AI racks do not behave like the CPU-heavy racks that shaped many existing data halls. NVIDIA describes hyperscale facilities moving from roughly 20 kW per rack to more than 135 kW per rack. Its GB200 NVL72 reference configuration reaches about 120 kW at rack scale. At those loads, forcing enough cold air through heat sinks becomes an expensive mechanical problem. Fans consume power, air paths compete for space, and a small recirculation failure can create fast thermal excursions.

Direct-to-chip liquid cooling attacks the bottleneck at the source. Cold plates collect heat from GPUs and CPUs, a coolant distribution unit transfers it to a facility loop, and the building rejects that heat outdoors. Liquid has far more heat capacity per unit volume than air, so the facility moves the same heat with smaller temperature differences and less bulk airflow.

AI data center cooling with direct to chip liquid cooled GPU rack
Generated infrastructure study: cold plates, manifolds, and rack-scale GPU trays shift heat removal from the aisle to the liquid loop.

This is why AI data center cooling now belongs in the capacity plan. A cluster cannot deliver tokens if the electrical service, network fabric, coolant distribution, or heat-rejection plant reaches its ceiling first. Buying accelerators without reserving thermal headroom can leave expensive hardware power-capped or underutilized.

AI data center cooling and the 45°C loop

A warmer supply temperature helps because heat rejection depends on the difference between the facility loop and outdoor air. Raise the loop temperature and dry coolers can work through more hours of the year. NVIDIA says a favorable-climate, dry-cooler design can run as a closed loop with near-zero cooling-water consumption, aside from limited chiller use in some conditions. Its public example contrasts conventional cooling-tower use of roughly 2.6 million gallons per MW each year with near zero for that design.

That claim describes a specific architecture and climate, not a universal guarantee. Still, the direction is clear. The same NVIDIA analysis notes that a 45°C inlet can leave at about 55°C after absorbing chip heat. That hotter return is useful: it improves the temperature available to a dry cooler and makes low-grade heat recovery more plausible where a nearby load exists.

The efficiency question is not simply “air versus liquid.” It is where the heat moves, at what temperature, and how many components sit in the path. A hybrid server may put cold plates on GPUs while retaining fans and air-cooled power or networking parts. NVIDIA describes Rubin as fully liquid cooled, including networking, which removes that residual fan dependency. A deployment team should ask exactly which components are on the liquid loop before applying full-liquid assumptions to a rack design.

NVIDIA also reports that GB200 NVL72 packs 72 Blackwell GPUs and 36 Grace CPUs in a rack-scale system. The company cites 25x better energy efficiency and 300x better water efficiency than traditional air-cooled architectures for its comparison. Those are vendor comparison figures, not a substitute for a site model. Use them as a starting hypothesis, then test against local weather data, utility tariffs, water constraints, and the chosen CDU and dry-cooler approach temperatures.

Dry coolers used for AI data center cooling beside an AI facility
Generated infrastructure study: high-temperature liquid loops can hand more of the heat-rejection job to outdoor dry coolers.

Why developers and platform teams should care

Thermal design changes application behavior through the capacity contract. Inference services see bursty arrival rates, long contexts, batching changes, model swaps, and failover events. A cluster sized only for average power may throttle when an agent workload pushes concurrent decoding or when traffic shifts after a node failure. Thermal headroom protects latency and throughput just as surely as spare GPU capacity does.

That means developers should expose the right signals to the infrastructure team. Track tokens per second, queue depth, GPU power draw, clock throttling, inlet and outlet coolant temperatures, and CDU differential pressure on the same dashboard. Correlate them during load tests. A rising queue with flat GPU utilization might be a software bottleneck. A rising queue with falling clocks and a narrowing thermal margin is a cooling problem.

Capacity models should also use useful work, not rack count. A practical metric is sustained tokens per second per delivered MW, measured at the facility boundary and under a realistic prompt mix. Include retries, networking, storage, and cooling power. That number makes a 120 kW rack meaningful. It also prevents a misleading comparison between a high-density system that runs steadily and a nominally similar system that must power-cap on hot afternoons.

A practical design checklist

  • Model the worst hour: use local dry-bulb and wet-bulb design conditions, not annual averages. Calculate dry-cooler capacity with the required approach temperature.
  • Separate loops: define the technology cooling loop, facility water loop, CDU heat exchanger, filtration, and glycol strategy. Confirm materials compatibility and water treatment.
  • Plan for failure: test pump, CDU, power-feed, sensor, and dry-cooler failures at a representative AI load. State the permitted time before thermal throttling begins.
  • Instrument the rack: capture supply and return temperatures, flow, pressure drop, leak detection, valve position, GPU clocks, and power caps. Alert on the rate of change, not only absolute temperature.
  • Specify service access: quick disconnects, dripless connections, isolation valves, and maintenance clearance determine whether a repair becomes a controlled job or a long outage.
  • Price water and energy together: compare PUE, water use effectiveness, compressor runtime, and peak-demand charges. A design with low fan power can still disappoint if the heat-rejection plant is wrong for the climate.

The 45°C milestone matters because it shifts cooling from a passive overhead to a design variable that can unlock compute density. NVIDIA’s 45°C liquid-cooling analysis explains the closed-loop and dry-cooler rationale, while its Blackwell efficiency overview documents the rack-density and water-efficiency claims used here. For teams building inference capacity, the practical conclusion is simple: validate thermal performance against the actual workload, weather file, and failure modes before the GPUs arrive. Run and scale models with Wiro.


Leave a Comment

Your email address will not be published. Required fields are marked *