top of page

The Cooling Challenge Behind the AI Data Center Boom

6 hours ago
5 min read

AI is changing the data center - and it is changing the way data centers need to be cooled.



The extraordinary growth in AI computing is driving much higher power and heat densities within data centers, and as a result, liquid cooling is moving rapidly from a specialist technology towards a mainstream part of data center infrastructure.


Recent industry coverage has focused increasingly on liquid cooling, coolant quality and the challenges of integrating new cooling technologies with existing infrastructure. At the same time, operators are looking at hybrid cooling, combining conventional air cooling with liquid cooling to support different generations of equipment within the same facility.


But there is another issue that deserves much more attention: How do you monitor an increasingly complex cooling environment effectively?


Modern data centers can contain a combination of:

  • Traditional air cooling

  • Direct-to-chip liquid cooling

  • Rear-door heat exchangers

  • Cooling distribution units (CDUs)

  • Chillers and heat exchangers

  • Pumps and secondary fluid loops

  • Temperature, pressure and flow sensors

  • Leak detection and coolant-quality monitoring


Each system produces operational data, and cooling is becoming a data problem.

The challenge is not necessarily generating more data. It is making sense of the data that already exists.


This becomes particularly important with hybrid cooling. An existing data center may have conventional servers that continue to rely on air cooling, alongside high-density AI infrastructure requiring liquid cooling.


The result is not simply an air-cooled facility becoming a liquid-cooled facility. It is a more complicated thermal environment in which multiple systems need to operate together.


Why monitoring matters


A cooling problem does not necessarily begin with a dramatic failure.

It might start with a gradual change in:


  • Temperature

  • Pressure

  • Flow

  • Pump performance

  • Coolant quality

  • Heat-exchanger efficiency

  • Energy consumption


Individually, these changes may appear insignificant. Viewed over time, however, they can provide valuable information about how a system is behaving.


Recent industry analysis has highlighted the growing importance of cooling-media quality in liquid-cooled data centers. Parameters such as conductivity, pH, turbidity and dissolved oxygen can provide information about the condition of cooling fluids and potentially help identify developing problems.


This illustrates an important shift in thinking.


Monitoring should not simply tell an operator when something has gone wrong. It should help reveal when something is beginning to behave differently.



From monitoring to analytics


There is an important distinction between collecting data and understanding it. A sensor can tell you that a temperature is 32°C. A monitoring system can tell you that the temperature has exceeded a predefined threshold.


Analytics can potentially provide something much more useful: Is this temperature behaving differently from the normal operating pattern for this part of the cooling system?


That requires looking at data in context and, importantly, looking at it over time.

As data centers become more complex, operators cannot realistically be expected to interpret every measurement manually. The volume of operational information is simply too great.


The opportunity is therefore to connect that information and make meaningful changes easier to identify.


Hybrid cooling makes visibility even more important


Hybrid cooling is likely to remain an important part of the transition to higher-density computing.


Existing facilities have substantial investments in conventional cooling infrastructure. At the same time, AI workloads can create thermal requirements that traditional air cooling alone may struggle to accommodate.


Rather than replacing everything, operators can combine technologies. But combining technologies also means combining data.


A cooling distribution unit, pump, heat exchanger, chiller and conventional HVAC system may all be contributing to the same overall thermal outcome.


Understanding each component individually is useful.


Understanding how they are behaving together is considerably more valuable.

This is why the data layer behind cooling infrastructure is becoming increasingly important.


The secondary cooling loop


The secondary fluid network deserves particular attention.


As liquid cooling becomes more widespread, cooling distribution units, pumps, pipework, heat exchangers and secondary fluid loops become an increasingly important part of the reliability chain. And all of these systems generate data.


Coolant condition can change gradually. Flow characteristics can change. Temperatures and pressures can drift. Equipment can become less efficient.


A single reading provides a snapshot.


A historical record can reveal a trend.


That distinction matters when the objective is to understand whether a cooling system is operating normally, gradually changing or developing an issue.


Where DataServe Analytics fits



This is where DataServe Analytics can provide value.


DataServe Analytics is designed to bring operational data together and provide a clearer view of what is happening across an organisation's infrastructure.


For cooling systems, that means moving beyond isolated readings and disconnected equipment.


Temperature, pressure, flow, energy and other operational information can be considered as part of the wider system, helping operators ask important questions:


What has changed?


When did it change?


Where did the change begin?


Is it an isolated event or part of a developing trend?



Is one part of the cooling infrastructure behaving differently from comparable equipment?

These are questions that become increasingly important as data-center cooling evolves.

Cooling is becoming mission-critical


The importance of cooling goes well beyond keeping equipment at an acceptable temperature.

It can affect:


Availability - thermal problems can ultimately affect computing performance and uptime.

Energy efficiency - inefficient cooling can increase the energy required to support the same IT workload.


Equipment reliability - persistent thermal or fluid-management issues can place additional stress on infrastructure.


Maintenance - identifying changes earlier can provide an opportunity to investigate before a problem becomes disruptive.


Capacity planning - as rack densities increase, understanding the actual performance of cooling infrastructure becomes increasingly important.


The industry is already beginning to explore increasingly intelligent approaches to cooling.


Recent work by Daikin and NTT DATA, for example, has examined using AI to predict server thermal conditions and coordinate HVAC, chiller and liquid-cooling systems rather than optimising each system independently.


That points towards an important development - the future of data-center cooling is not simply about installing more powerful cooling equipment.


It is about creating a data-driven thermal management environment.


The future is better visibility.


The AI data-center build-out is creating enormous demand for computing power, electricity and cooling.


Liquid cooling is becoming increasingly important. Hybrid architectures offer a way of combining existing infrastructure with new high-density systems. Cooling systems themselves are becoming more sophisticated, while the quality and condition of cooling fluids and supporting infrastructure are receiving greater attention.


All of this creates more operational data.


The challenge is turning that data into useful information.


Because when cooling becomes mission-critical, knowing what is happening is no longer optional.


DataServe Analytics provides the data and analytics layer needed to help organisations gain a clearer understanding of their operational infrastructure - helping turn an ever-growing volume of data into meaningful visibility.


The next generation of data centers will need more than better cooling. They will need better insight into the cooling they already have.



Comments


bottom of page