Liquid Cooling and AI: Wait, watch or act?

Two workers inside a data centre

Whether they're looking to upgrade an existing data centre to run AI systems or are planning infrastructure for a local AI research project, data centre operators worldwide face a key strategic question: should we wait for the next generation of GPUs, or should we get started now with the hardware currently available? This decision, which seems simple on the surface, has far-reaching consequences, and the answer can vary by region. While the Asian market tends to take a pragmatic approach, American hyperscalers often set key infrastructure standards, and Europe is creating a common foundation through regulation and ambitious sustainability goals. But a closer look reveals that these different approaches aren't necessarily a disadvantage; rather, when combined, they can help build important expertise. The question of the right timing is ultimately a matter of experience and practice, and those who don't start early enough may find themselves left behind by upcoming AI technologies.

 

STULZ AI technologies

 

The availability problem: why waiting can become a trap

Some data centre operators might think: why invest in the current generation of chips now, when a better and much more efficient generation of GPUs will be released in a few months anyway? However, those who follow this logic may be overlooking a fundamental market problem. New generations of GPUs and TPUs are introduced, especially at the beginning, with clear capacity limits. Major investors, such as leading cloud providers and hyperscalers, then snap up the bulk of the initial allocations. Companies that wait and hope to jump into an AI project right after the first wave of shipments may be in for a surprise, as there could be shortages again even for the next generation of chips. Thus, the vicious cycle repeats itself, and data centre operators who are still hesitating today risk finding themselves in an unfavourable market position right from the start. While competitors' first projects are already running on the latest AI hardware, companies that have been too hesitant will then struggle with capacity availability, precisely for the generation they've been waiting for so long.

Building expertise as a true competitive advantage

The decisive argument for an early entry into the field of AI infrastructure is therefore not only technical but also strategic in nature. Building expertise as early as possible is crucial. Those who start now with NVIDIA Blackwell or other current systems can develop genuine competitive capabilities and build technical understanding of how AI chips react under load in practice, what thermal fluctuations occur, and how responsive cooling systems must be designed. These insights can then be transferred much more easily to the next generation, because operators will know what to look out for. They will master the operational management of liquid-cooling systems, the maintenance of CDUs, and the control of temperature profiles. This expertise is essential, as it enables operators to build long-term supplier relationships, understand scaling mechanisms, and know how to implement "pay-as-you-grow" models in the AI sector.

 

STULZ Liquid cooling map

 

Three regional strategies, one global goal

The global market for AI infrastructure currently exhibits three very different strategies. All of these approaches incorporate important perspectives and elements for the sustainable success of projects.

 

STULZ Asia map

 

The Asian market, for example, leads the way in terms of speed of adaptation. In Asia, decisions are often made within a few weeks, rather than over several months. Pilot projects systematically serve as test labs. Data centre operators start on an experimental basis, optimise operations as they go, and document their experiences. Mistakes are welcomed, as they present opportunities to build expertise. This mindset has ultimately given Asia a valuable head start. There are already proven reference projects, tested modular scaling strategies, and local adaptation options tailored to specific climatic conditions. Each project thus contributes to the collective body of knowledge.

 

STULZ USA Map

 

The American market, on the other hand, is dominated by large hyperscalers. Their business model is based not on experimentation, but on rapid scaling. Amazon, Google, Microsoft, and other market leaders are investing simultaneously in various solution tiers, ranging from standard systems to fully customised AI infrastructures. This approach establishes global standards and provides regulatory stability for long-term planning. However, the downside of this approach is also clear: only large players truly benefit from this high-volume market.

 

STULZ Europe map

 

While Europe has been more cautious in adopting AI, it leads the way in sustainability standards and key guidelines. The Energy Efficiency Act (EnEfG) requires new data centres to comply with a maximum PUE of 1.2 starting in 2026. These requirements effectively make liquid cooling the de facto minimum standard. The EU Energy Efficiency Directive also requires heat recovery analyses for new facilities exceeding 1 MW. This complexity inevitably leads to longer planning cycles, but also to better-designed, more sustainable infrastructures. The European approach favours proven technologies and thus has a competitive advantage that should not be underestimated: ESG compliance.

Another aspect often overlooked in investment decisions is the global supply chain situation. Most AI accelerator chips are manufactured in Asia. But it's not just chips. GPU packaging, cooling plates, and flow-control components also often come from Asia. This global dependence poses a real risk. Geopolitical tensions could trigger export restrictions. Those who wait today are not only waiting for better hardware, but also for supply chains that could quickly grind to a halt in the event of geopolitical shifts.

Depending on how you look at it, however, the apparent fragmentation of the global IT market can also be seen as a strength. Asia's flexibility quickly generates expertise. The US successfully scales key architectural decisions for the mass market. Europe standardises them through regulation. This division of labour accelerates the global availability of liquid-cooling solutions. Data centre operators worldwide benefit from this. They can access tested reference designs from Asia, proven scaling models from the US, and regulatory-compliant standards from Europe. Those who understand this global accumulation of expertise do not adopt a regionally limited perspective but rather lay the foundation for well-thought-out, strategically balanced investments.

Modular planning instead of rigid investments

Given the global market dynamics surrounding the adoption of AI solutions, the immediate launch of AI pilot projects is the logical next step, especially for European and American data centre operators. In this phase, the goal is to establish essential benchmarks for operating temperatures, energy efficiency, and availability, which can then serve as the foundation for all subsequent expansion phases. Expertise is built in three key areas: operational activities and supplier evaluations yield insights that can be directly applied, while operators explore the practical limits of their infrastructure during the scaling phase. To reduce complexity costs, it is also advisable to rely on co-engineering processes with selected technology partners rather than traditional bidding practices.

The infrastructure should also be designed for rack capacities of approximately 125 to 140 kW, but should be able to scale up by one or two generations if necessary, without requiring a complete overhaul of the basic infrastructure, such as refrigerant lines, power supply, or ventilation systems. This can be achieved, among other things, by including pre-dimensioned hydraulic paths in the initial design, which are only put into operation later as the AI system's utilisation grows. A scalable CDU architecture enables modular expandability. Predefined interfaces for future generations, as well as a digital twin established during the planning phase, open up opportunities for optimisation later on. With this investment protection strategy, data centre operators do not pay today for capacity that will not be used tomorrow. They plan modularly so that expansions are possible even without costly redesigns.

Frequently asked questions

Should we wait for the next generation of GPUs before investing in liquid cooling?

Waiting rarely pays off. New GPU and TPU generations launch with tight capacity limits, and hyperscalers and large cloud providers absorb most of the initial allocation. Operators who hold off often find the same shortage repeating with the following generation, while competitors are already running live AI workloads and building operational knowledge.

Why do AI workloads need liquid cooling rather than air cooling?

AI accelerators generate heat loads and thermal fluctuations that air cooling struggles to handle at modern rack densities. Liquid cooling removes heat far closer to the source, which is why new AI infrastructure is increasingly designed around it from the outset rather than retrofitted later.

What rack density should a new AI data centre be designed for?

Plan for roughly 125 to 140 kW per rack, with headroom to scale one or two generations further. The important part is that scaling up should not require replacing the basic infrastructure, so refrigerant lines, power supply and ventilation need to be sized with that growth in mind.

What is a CDU, and why does its architecture matter?

A coolant distribution unit manages the flow, pressure and temperature of liquid between the facility system and the IT equipment. A scalable CDU architecture matters because it allows capacity to be added in stages, so operators expand as demand grows instead of paying upfront for capacity that sits idle.

How is regulation pushing data centres towards liquid cooling?

Efficiency regulation is tightening. Germany's Energy Efficiency Act requires new data centres to meet a maximum PUE of 1.2 from 2026, which effectively makes liquid cooling the minimum viable standard, and the EU Energy Efficiency Directive requires heat recovery analyses for new facilities above 1 MW. These rules lengthen planning cycles but produce better-designed, more sustainable facilities.

How can we invest in liquid cooling without overcommitting?

Through modular planning. Start with a pilot project to establish benchmarks for operating temperature, energy efficiency and availability, include pre-dimensioned hydraulic paths that are only commissioned as utilisation grows, and define interfaces for future generations at the design stage. A digital twin built during planning makes later optimisation easier. This way expansion is possible without a costly redesign.

About the author - Aaron Jürgens

Jonathan Aaron Jürgens is a Team Lead Liquid Cooling Global Product at STULZ GmbH, where he has been shaping the future of data centre cooling since 2020. Holding a bachelor's degree in industrial engineering, Jonathan plays a key role in the development and market implementation of the CyberCool CDU. With hands-on experience from over 100 global liquid cooling projects since the CDU's launch, Jonathan has established himself as the in-house expert on liquid cooling infrastructures. To further share this expertise, he conducts liquid cooling seminars worldwide, equipping data centre stakeholders with the knowledge needed to optimise cooling performance while enhancing sustainability.

Jonathan Aaron Jürgens