Direct-to-chip cooling failure causes week-long data center outage and revenue loss
A data center experienced a week of downtime and significant revenue loss after a contractor used municipal water instead of deionised water in a direct-to-chip cooling loop. The incident highlights the critical risks of improper coolant chemistry and the necessity for continuous monitoring in high-density AI infrastructure.
Key Takeaways
- Contractor error using municipal water instead of deionised water caused calcium phosphate scale to plug cooling flow
- Ecolab manager Kelsey Nowicki reports that problems in high-density loops surface in weeks rather than months
- Frequent hardware change-outs in AI infrastructure increase risks of introducing contaminants and dissolved oxygen
- Continuous monitoring is recommended over periodic sampling due to narrow operating margins for propylene glycol concentration
Why It Matters
The shift toward high-density AI infrastructure requires a transition from viewing cooling as standard plumbing to treating it as a controlled process system. For streaming platforms relying on these data centers, even minor chemical imbalances in direct-to-chip loops can lead to rapid hardware failure and service interruptions. As operators scale up liquid cooling to handle increased compute loads, the margin for error in maintenance and commissioning narrows significantly. Watch for data center operators to mandate automated chemical monitoring systems as a standard requirement in service-level agreements to mitigate these specific infrastructure risks.
Read full article at capacityglobal.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source