From a Power Transient to a 15-Hour Outage: Lessons from Google’s Netherlands Data Center Cooling Incident
Date Published

From a Power Transient to a 15-Hour Outage: Lessons from Google’s Netherlands Data Center Incident
Introduction
The incident report published on July 15, 2026 identified a 3-millisecond utility voltage dip as the initiating event.
The transient and the operation of protective circuit breakers affected both utility feeds.
One side successfully transferred to DRUPS power, but the chilled-water circulation pumps did not restart automatically.
Rising temperatures led to the shutdown of IT equipment, while the disruption to three services lasted a total of 14 hours and 55 minutes.

According to the report published by Google on July 15, 2026, an extremely brief electrical grid event at the Netherlands data center serving the Google Cloud europe-west4-a zone developed into a service outage lasting nearly 15 hours. The case is particularly important because it was not simply about a loss of power. It exposed the dependencies between the power transient, protective operation, backup power, and the restart of the cooling system.
Two independent utility feeds do not by themselves provide end-to-end resilience. Cooling auxiliaries, controls, and restart logic must also remain operational throughout the event.
Who is this analysis for?
This article is intended for data center operators, facility managers, IT leaders responsible for critical infrastructure, designers, investors, and technical procurement teams. Its professional focus is not the software layer of the cloud service, but the coordination of the electrical and mechanical infrastructure supporting it.
The 3-millisecond transient was brief, but the response of the complete power and cooling chain determined its consequences.
The industry significance of the incident lies in demonstrating that redundant power feeds are not equivalent to resilience across the complete technology chain. Availability also depends on which auxiliary systems remain operational after a transient, which systems stop, and how safe automatic recovery is performed.
What happened at the Netherlands data center?
Realistic integrated testing verifies not only power transfer but also the automatic recovery of pumps, controls, and thermal protection processes.
According to the published summary, a 3-millisecond utility voltage dip and the operation of protective circuit breakers affected both utility feeds. One side of the power system successfully transferred to a DRUPS system. A DRUPS is a dynamic, rotating uninterruptible power supply solution designed to bridge critical loads and support power continuity during utility disturbances.
However, the transfer alone did not restore the facility to its complete operating state. The chilled-water circulation pumps did not restart automatically, resulting in a loss of cooling capacity and rising temperatures in the data hall. Servers, storage systems, and network equipment were shut down to protect the hardware. According to the report, three services experienced a total disruption of 14 hours and 55 minutes.
Why is the complete supply chain more important than the number of power feeds?
Data center redundancy is often described by the number of power sources, UPS systems, generators, or utility feeds. Actual availability, however, must be assessed from end to end. It is not sufficient for the critical IT load to continue receiving power through one supply path if the pumps, controllers, or control circuits required for heat removal do not recover.
The cooling system can itself consist of several interdependent layers: heat-generating IT equipment, room-side heat transfer, fluid circulation, heat rejection, electrical auxiliaries, automation, and monitoring. The failure of a single component without a functioning alternative or suitable restart logic can restrict the capacity of the entire system.
The incident should therefore not be interpreted solely as a cooling problem. It was a system-level issue involving power supply, automation, and recovery.
What does this mean from an operational perspective?
The difference between the brief transient and the prolonged service impact indicates that operational readiness should not be optimized exclusively for surviving the initial failure. Recovery time is at least as important.
After a utility disturbance, operators need a clear understanding of:
which mechanical loads stop and which remain operational;
which equipment restarts automatically;
which conditions or interlocks may prevent a restart;
the sequence in which cooling capacity is restored;
how much thermal ride-through time is available;
which alarms require immediate human intervention;
how the IT load will be reduced in a controlled manner if cooling is not restored.
Thermal reserves are not unlimited. The room, fluid circuits, and structural mass may slow the temperature increase for a period, but the actual time window depends on the load, system design, and current operating state. Recovery procedures should therefore be based on facility-specific data.
A common failure mode: testing components without an integrated system test
A common risk is that the UPS, generator, switchgear, and cooling units are tested successfully as individual components, while the complete event sequence is not examined under realistic conditions. A successful DRUPS transfer, for example, does not automatically prove that cooling pumps, variable-frequency drives, controls, and communication links will also reach their intended operating states.
Short utility disturbances require particular attention. They may not cause a sustained power outage, yet they can still trigger protective devices or changes in control states. Testing should therefore address transients, transfer states, and failed automatic restart conditions.
There can be a material difference between documented redundancy and redundancy proven under operating conditions.
What should be reviewed from a technical perspective?
Based on the incident, it is advisable to trace the power path of every critical load in the cooling chain. The review should determine which distribution branches supply pumps, control panels, automation units, valves, sensors, and network devices, as well as what happens to each component during a voltage dip or circuit-breaker operation.
Key areas for review include:
1. Protection coordination: Do circuit breakers and protective devices operate in a sequence consistent with the intended selectivity? 2. Auxiliary power supply: Are the loads essential to cooling connected to appropriately protected power paths? 3. Restart logic: Do systems return to their operating state automatically and safely after a brief transient? 4. Interlocks and default states: Could a lost signal or interrupted communication link unnecessarily prevent recovery? 5. Monitoring: Can operators clearly identify a discrepancy between the electrical state and actual cooling performance? 6. Emergency procedures: Is it defined when loads must be reduced or equipment shut down in a controlled manner?
Recommended next step
The first recommended step is a joint review of electrical, mechanical, and automation dependencies. This should cover not only single-line diagrams but also control power, communication links, restart conditions, and operating procedures.
A risk-based integrated testing program can follow. Tests should be designed to avoid placing the live environment at unnecessary risk. Phased execution, temporary backup capacity, simulation, or testing under limited load may be appropriate. The objective is not a dramatic full shutdown, but controlled verification of critical transitional states.
The findings should be incorporated into the alarm matrix, maintenance plan, operator training, and recovery procedures. Repeat validation is required after modifications have been implemented.
The Digital Technologies engineering perspective
The resilience of critical infrastructure can only be evaluated by examining electrical, mechanical, and IT interfaces together. During an audit or modernization project, it is therefore not sufficient to inspect the UPS and cooling equipment separately. Power paths, auxiliary systems, protection settings, controls, and recovery processes must be treated as one integrated system.
Digital Technologies supports the design, review, modernization, commissioning, and maintenance of critical power and data center cooling systems through a vendor-independent approach. In operating facilities, phased execution, documentation of temporary operating states, and continuous control of service risk are particularly important.
Conclusion
The incident at Google’s Netherlands data center demonstrates that the consequences of a 3-millisecond electrical event can last far longer than the transient itself. The successful transfer of one power path to DRUPS did not prevent the loss of cooling capacity because the circulation pumps did not restart automatically.
The practical lesson is clear: redundancy should not be treated as an equipment list, but as an operational, tested, and recoverable system. During the next audit or integrated test, the question should not only be whether a backup power source is available. It should also be whether the complete cooling chain actually returns to the required operating state after power has been restored.
Sources
Google Cloud incident report: https://status.cloud.google.com/incidents/3BvH3LVGcupoYqV6F4Nw
Technical analysis by The Register: https://www.theregister.com/off-prem/2026/07/21/google-cloud-outage-shows-its-still-hard-to-understand-hyperscalers-real-resilience-regimes/5275405
Let's start work together

Locations in Germany & Hungary
Related Services
Additional related links and relevant content in the same topic area.






