The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount

Appendix F

Failure-Mode / FMEA Catalog

Random failures and adversarial scenarios need distinct initiating evidence and controls; link them only where a validated topology shows that they converge on a shared top event or protective response.

What you'll decide here

  1. Use this appendix as a lookup, not a narrative: find your failure mode in the master table, then jump to the owning chapter (last column) for the engineering derivation — this catalog summarizes, it does not replace, the canonical treatment.
  2. Keep random and adversarial initiators, likelihoods, authorities, observability, persistence, controls, and recovery evidence distinct; cross-link only the validated shared top event or protection function.
  3. Treat the blast-radius column as the input to your fault-domain and RBD work (Chapter 12.1 / 12.5): a mode that strands one rack and a mode that trips the campus POI sit in different availability tiers and deserve different redundancy spend.
  4. Walk the propagation column for cascade interlocks. Most catastrophic outages here are not the first fault — they are a single fault that defeated a shared mitigation (one CDU, one fuel header, one protection setting) and took an entire fault domain with it.
  5. Pair detection latency against propagation speed for each mode. Where the fault propagates faster than your detection-plus-actuation loop (load-step→grid trip, HBM runaway), the only viable mitigation is preventive or inertial, never reactive — design accordingly.

This appendix consolidates the failure modes scattered across the engineering chapters into one uniform FMEA register. It is referenced from the resilience-standards chapter (Chapter 12.1), the reliability rethink (Chapter 12.2), the component-failure-rate chapter (Chapter 14.3), and the integrated-systems-test chapter (Chapter 13.6); and it is consumed directly by the quantitative availability model in Chapter 12.5, which draws its top events and common-cause couplings from the rows below. The canonical engineering of each mode lives in the chapter named in the right-most column — this catalog is the index and the cross-walk, not the derivation.

Some random and adversarial scenarios can converge on a similar top event or protective response, but they do not automatically share one record. Keep initiating evidence, likelihood, authority, observability, persistence, propagation, security controls, and recovery distinct. Where the validated topology joins the paths, link the records to the shared top event and reuse the proven protective response; do not claim every random failure has an attacker twin. The cyber-physical analysis is in Chapter 11.10.

How to read a row

Each failure mode is recorded against six fields, applied identically across every table so the catalog is sortable and comparable:

  • Trigger — the initiating event (random cause first; the induced/attacker path is noted where it differs materially).
  • Propagation path — how the fault spreads, and critically which shared mitigation it defeats to escalate from a local fault to a cascade.
  • Detection — the sensing modality and its characteristic latency relative to propagation speed (the decisive ratio).
  • Blast radius — the fault domain affected at full propagation: node, rack, row, hall, or campus/POI.
  • Mitigation — the preventive or containing control, classified as preventive (stops the trigger), inertial (buys ride-through time), or reactive (acts after detection).
  • Recovery — the path back to service and its characteristic time-to-restore (TTR).

Where propagation outruns detection-plus-actuation, the only effective controls are preventive or inertial — which is why that ratio is called out per mode in the notes.

Master FMEA catalog — thermal & mechanical (cooling) modes

Cooling-system failure modes
Failure modeTriggerPropagation pathDetectionBlast radiusMitigationRecoveryOwner
Coolant-leak cascadeQD/manifold/cold-plate breach, hose chafe, gasket creep; induced: spoofed dew-point setpoint forcing condensationLocal drip → conductive coolant on busbar/PDU → arc/short → de-rate or trip of the powered branch; if it reaches a shared CDU controller, the whole CDU loop and its row drop togetherFloor/leak-rope sensors + CDU flow/pressure-decay; latency seconds-to-minutes; propagation can outrun it on a high-flow breachRack → row (if the CDU is shared); hall if isolation valves are absentNegative-pressure loops (leak draws air in, not coolant out); dripless UQDs; zoned isolation valves; per-rack leak detection; N+1 CDU with independent controllers (preventive + reactive)Isolate the branch, drain/flush the loop, replace the failed coupling, re-pressure-test, re-fill, re-commission worst-case branch; TTR hours per rack5.11
CDU / pump failurePump bearing/VFD failure, seal loss, filter blockage, control-board fault; induced: malicious VFD firmware or controller DoSFlow may be maintained, reduced, redistributed, or lost depending on pump arrangement, check valves, control sequence, loop volume and thermal mass, standby readiness, and IT throttling; temperature response follows that project topologyPump state, differential pressure, flow, coolant temperatures, valve position and IT telemetry; required detection and trip timing come from the validated transientThe affected hydraulic zone — rack, row, or larger only where the topology shares the failed resourceQualified duty/standby or N+1 pumps and CDUs, independent power/control where required, validated switchover, adequate thermal/flow ride-through, and coordinated IT throttlingExecute the validated failover or controlled IT response, isolate and repair the failed train, then re-qualify flow and controls; recovery time is topology- and failure-specific5.11
Cooling-controls transient excursionSynchronized GPU load drop (job ends / checkpoint pause) → loop heat input collapses faster than valves/VFDs can slew; setpoint hunt / control-loop oscillationOn a rapid load drop, supply-coolant temp overshoots downward → transient dew-point excursion → condensation risk on cold surfaces; or anti-hunting failure drives sustained oscillation that fatigues actuators and destabilizes neighboring loopsCoolant supply-temp rate-of-change, dew-point margin sensor, valve-position hunting; detectable but the excursion window is briefRow → hall (controls coupling); condensation risk is local to cold surfacesSlew-rate limits on control valves and pump VFDs; anti-hunting tuning; dew-point margin floor; coordinate cooling setpoints with the rack BBU/BESS load-smoothing spine (preventive)Re-tune control loops, restore dew-point margin, dry/inspect any condensation; TTR minutes, no hardware loss if caught5.12
HBM thermal runawayCold-plate contact loss, TIM pump-out, local flow starvation, or sustained over-temp on a stacked-DRAM site; induced: CDU disablement holding flow at zeroHBM junction temp climbs → ECC error rate rises → uncorrectable error / package damage; on a tightly-coupled training step the failed device stalls the synchronous collective and the whole job stalls behind the stragglerPer-die thermal telemetry, ECC/CE rate trend, GPU throttle flags; trend-detectable early, but runaway is fast once contact is lostNode (the GPU/HBM package) → job (synchronous training stalls on the straggler)Thermal screening/burn-in pre-deployment; ECC-rate alarming with proactive drain; flow-failure throttle floor; hot-spare nodes so the scheduler evicts and replaces the straggler (preventive + reactive)Evict the node, fail the job over to a hot spare, RMA the package; training resumes from last checkpoint (TTR = checkpoint interval + restart)14.3
Coolant chemistry/flow envelope per ASHRAE TC 9.9 (5th ed.) liquid-cooling guidelines and OCP Liquid Cooling white papers; CDU/QD practice per Vertiv/nVent/Equinix. Owning chapters in last column. Dual-use induced path per Chapter 11.10.

A direct-to-chip loop's response to pump or CDU loss is topology-, control-, and inertia-dependent. Pump arrangement, check valves, loop volume and metal mass, accumulator behavior, power ride-through, valve sequence, standby readiness, rack thermal mass, and IT throttling determine whether flow is maintained or lost and how quickly junction temperatures rise. Detection, protective action, and recovery timing must therefore come from the validated project transient, not from a universal seconds-versus-minutes rule. Engineer and test the complete cause-and-effect in Chapter 5.11; feed the measured response into the reliability model in Chapter 12.2.

Master FMEA catalog — electrical & power modes

Power-chain & grid-interface failure modes
Failure modeTriggerPropagation pathDetectionBlast radiusMitigationRecoveryOwner
Simultaneous-GPU-load-step grid tripThousands of GPUs ramp in lockstep at job start/stop/checkpoint; di/dt event on every step; induced: malicious power-cap firmware forcing a synchronized stepAggregate ramp >1,000 MW/s presented to the POI → voltage/frequency disturbance → if the load-smoothing spine is absent, upstream protection or generators see a step they cannot follow → tripPower-quality metering at the POI, PMU/PQM; fast, but the di/dt event is faster than any reactive control — millisecondsCampus (POI) → contributes to wide-area grid disturbanceThe chip→BBUBESS smoothing spine (on-package capacitance → rack BBU → facility BESS); software ramp-rate limits and regulated wind-downs; grid-forming inverters (preventive + inertial)Re-energize per utility ride-through procedure; restore smoothing controls; no hardware loss if the spine held; TTR minutes if ride-through succeeded4.5
Utility ride-through / voltage-disturbance eventGrid fault (e.g. 230 kV line fault) causes a voltage sag at the POI; sensitive customer-side protection drops the loadUndervoltage response across facility loads → cumulative ~1,500 MW customer-side load reduction during the July 2024 six-fault NoVA sequence → the reduction itself destabilizes the grid, a self-reinforcing reliability problem NERC flagged at Level 3POI relays, undervoltage/under-frequency elements, PMU; the disturbance is sub-cycle to cyclesCampus (full load drop) → wide-area gridFault-ride-through settings tuned to stay online through the sag (SEL/relay, UPS, undervoltage-load-retention); reactive/voltage support toward the POI; ride-through posture aligned to the NERC 2026 Level-3 alert's recommended large-load envelope — guidance, not penalty-backed; TPL-001/PRC standards bind transmission planners and generators, not data-center loads (preventive)Auto-recover as the sag clears if ride-through held; if tripped, sequenced re-energization and load ramp; TTR minutes4.10
BESS thermal runawayCell defect, overcharge, internal short, or cooling loss in an LFP facility battery; induced: BMS spoof disabling cell balancing/thermal protectionSingle cell vents → exothermic chain to adjacent cells → module-level runaway → fire/off-gas if pack-level isolation and venting fail; loss of the BESS also removes the ride-through and load-smoothing it was providingCell voltage/temp telemetry, off-gas (H2/CO) detection, BMS fault flags; early-warning gas detection precedes thermal runaway by a useful marginBESS enclosure → adjacent enclosures/room if propagation isolation failsLFP chemistry (higher thermal-runaway threshold than NMC); module-level thermal isolation and dedicated venting/deflagration paths; off-gas detection with pre-emptive isolation; physical separation of BESS from IT (preventive + reactive)Execute the pre-incident response plan: isolate electrically, keep ventilation/deflagration paths clear, monitor off-gas, coordinate fire-service response on the UL 9540A / NFPA 855 basis of design — never open or manually suppress the enclosure ad hoc; replace module/pack; re-commission; TTR hours-to-days; ride-through reverts to UPS/BBU meanwhile4.5
Fuel-supply interruption (on-site generation)Firm pipeline curtailment (correlated cold-snap), valve/compressor failure, or fuel-quality (Wobbe/dew-point) excursion; LNG/CNG storage depletionLoss of fuel → on-site turbines/engines de-rate or trip → if the site is islanded or grid-import is constrained, generation cannot meet IT load → controlled load-shed or outageFuel header pressure, Wobbe-index/dew-point analyzers, tank level, generator load; minutes of warning on slow depletion, immediate on a hard cutCampus (islanded sites) → partial if grid-import backstops'Synthetic-firm' fuel structure (multiple pipelines + interruptible + on-site LNG/CNG storage); dual-fuel switching; fuel conditioning to spec; sized on-site storage for correlated-curtailment duration (preventive)Switch fuel source / draw down on-site storage, restore generation; coordinate curtailment with curtailable-load agreement; TTR depends on storage sizing vs outage duration4.9
Load-step magnitudes: synchronized GPU draw can swing 30%→100% in milliseconds, aggregating to >1,000 MW/s at GW scale (NVIDIA/Microsoft/OpenAI joint findings, 2025). The cumulative ~1,500 MW customer-side load reduction across a six-fault, 82-second 230 kV sequence is the NERC Level-3 alert motivating case. Owners in last column.
30%→100% in ms
synchronized GPU power swing per load step; aggregates to >1,000 MW/s ramp at GW scale
~1,500 MW
data-center load loss during the six-fault, 82-second July 2024 VA 230 kV sequence — the NERC Level-3 alert motivating case
1 per ~3 hr
mean interruption rate on Meta's 16,384-GPU H100 Llama 3 cluster; 466 interruptions (419 unexpected) over a 54-day window
30.1% / 17.2%
share of training interruptions from faulty GPUs / HBM3 memory; network switch+cable 8.4%; >90% effective training time maintained
45 °C maximum liquid inlet; 65 °C maximum liquid return (separate limits)
QCT GB200 NVL72 QoolRack reference maxima: 45 °C liquid inlet and 65 °C liquid return; separate limits, not a selected operating pair
~1.8–1.9 L/kWh
industry-avg evaporative WUE (best-in-class 0.3–0.7; Microsoft FY2025 0.27)

Master FMEA catalog — connectivity, compute & data-integrity modes

Network, optics, compute & data-integrity failure modes
Failure modeTriggerPropagation pathDetectionBlast radiusMitigationRecoveryOwner
Fiber cut (inter-/intra-DC)Backhoe/construction strike, conduit failure, or DCI route cut; induced: deliberate cut of an un-diverse routeLoss of a fiber path → if scale-across/DCI is single-routed, the affected campus or training partition is severed → distributed training stalls or a metro site loses connectivityOptical LOS/LOF alarms, OTDR, BER collapse; immediate at the physical layerLink → partition/campus (distributed training) or metro site reach (inference)Physically diverse, geographically separated fiber routes; protected DCI (ZR/ZR+ with restoration); for training, topology that degrades gracefully on a partition loss (preventive + reactive)Restore over the diverse path automatically; physical splice repair on the cut route (TTR hours-to-days for the splice, seconds for protected failover)3.6
Optics flap stormMarginal transceiver, dirty/over-bent connector, thermal cycling, or firmware bug causing repeated link up/down; induced: thermal attack on the optics environmentOne flapping link → routing reconvergence churn / ECMP rehashing → packet loss and tail-latency spikes propagate across the fabric → on a tightly-coupled collective, the flapping link gates the whole all-reduce and tanks MFUPer-port link-flap counters, BER/FEC-error trend, CRC errors; trend-detectable but a storm builds in seconds-to-minutesLink → fabric pod (reconvergence churn) → job (collective stalls)Pre-install optics burn-in / BER screening off the critical path; FEC margin headroom; auto-quarantine of flapping ports; CPO to remove pluggable failure points where deployed (preventive + reactive)Quarantine and replace the offending transceiver; let routing reconverge; TTR minutes to swap, plus reconvergence8.9
SDC corruption event (silent data corruption)Marginal/defective silicon (a 'mercurial core'), aging, voltage/thermal margin loss producing a wrong result with no error flag; induced: fault-injection on a known-marginal deviceA miscomputed value flows silently into gradients/activations → corrupts the model state or an inference result → undetected for hours-to-days, potentially poisoning a checkpoint and forcing a roll-back of all work since the last clean checkpointNo native hardware flag — requires a dedicated detection program (Fleetscanner periodic, Ripple in-fleet, Hardware Sentinel runtime); latency hours unless instrumentedNode (the mercurial core) → job/model (silent corruption of state) → potentially every downstream consumer of a poisoned checkpointProactive SDC-hunting at scale; redundant/checksummed compute on critical paths; PVF-aware placement; roll-back to a verified-clean checkpoint; quarantine the device (preventive + reactive)Identify and quarantine the mercurial core, roll back to last verified-clean checkpoint, re-run; TTR = work since clean checkpoint (the reason hyper-frequent checkpointing pays off)14.3
Water-curtailment eventDrought-stage ordinance, basin/withdrawal-permit curtailment, or reclaimed-water supply interruption on an evaporative-cooled site (grid-curtailment orders like ERCOT's SB6 regime are an electrical event, not a water one)Loss/limit of make-up water → evaporative/cooling-tower capacity drops → heat-rejection ceiling falls below IT load → forced de-rate of compute or a curtailment-driven load-shedMake-up water flow/level, WUE telemetry, basin level, curtailment-order receipt; advance notice on scheduled curtailment, immediate on a hard orderHall → campus (heat-rejection limited)Closed-loop / zero-evaporation cooling design (designs water out of the risk); reclaimed/non-potable sourcing; on-site water storage; curtailment-tolerant workload scheduling (batch defers); thermal-storage buffer (preventive)Shift heat rejection to the closed-loop/dry path or draw down water storage; defer curtailable batch load; TTR = curtailment duration; structural fix is closed-loop conversion3.7
SDC detection program per Meta (Fleetscanner/Ripple/Hardware Sentinel) and OCP SDC-in-AI white paper; optics reliability per SemiAnalysis/NVIDIA CPO analyses; fiber/latency per Chapter 3.6. Owners in last column.

Three of these four modes share the opposite signature from the cooling and load-step modes: a long detection latency against a slow-burning blast radius. An SDC event can poison a checkpoint and sit undetected for days; an optics flap can quietly erode MFU before anyone correlates the tail-latency to a single marginal transceiver; a fiber cut on a poorly-instrumented diverse path can go unnoticed until the protect path also fails. For this class the decisive investment is detection, not faster actuation: a dedicated SDC-hunting program, per-port flap/FEC telemetry, and verified-clean checkpointing convert an invisible, unbounded-blast-radius fault into a bounded, recoverable one. The empirical fault taxonomy and the AFRs that feed these rows into the availability model are the canonical content of Chapter 14.3; the SDC detection program is treated there and demonstrated at commissioning in Chapter 13.6.

Common-cause couplings & cascade interlocks

For the availability model, the catalog's most useful output is the couplings between modes — the shared resources whose failure makes two independent-looking modes fail together. These are the common-cause terms an RBD or fault tree must capture, or it will badly over-state availability. The table below names the cross-mode interlocks worth modeling explicitly.

Candidate cross-mode couplings — validate before modeling a shared event
Candidate shared conditionModes it may coupleCondition to verifyDesign / model response
CDU, controller, loop, or power source shared by a hydraulic zoneLeak escalation, loss of flow, thermal protectionWhich racks actually share flow, control, isolation and power; whether one fault defeats standbyModel one shared event only for the confirmed fault domain; separate or isolate trains where required
One storage system used for smoothing and ride-throughLoad-step response, ride-through, storage isolationWhether the same cells, inverter, controls, protection or bus provide both functionsSeparate functions or model the confirmed common dependency and its protection states
Fuel source or delivery path exposed to one regional eventFuel interruption and generation availabilityContract, pipeline, storage and transport independence during the named event durationRepresent the named supply event explicitly; do not infer independence from two contracts
Fiber routes sharing conduit, crossing, right-of-way, building entry or carrier equipmentRoute cut and loss of protect pathSurveyed physical diversity end to end, including meet-me and power dependenciesUse an explicit route-loss event where sharing is confirmed; correct false diversity
Checkpoint policy and last-known-good stateRestart and rollback recovery timesIntegrity validation, cadence, write path and corruption horizonTreat as a recovery-time dependency, not a hardware common-cause beta
Shared BMS/DCIM/SCADA authority or communicationsCooling setpoint, flow and load-control excursionsWhich commands, sensors, interlocks and overrides share one control planeModel only confirmed shared authority; provide independent protection where the hazard basis requires
Each row is a hypothesis to test against the project's topology, controls, physical routes, operating rules, and event data. Confirmed coupling may become an explicit shared event; it is never automatically a numeric beta-factor.

Using this catalog in the availability model

The rows above are inputs to reliability work, not ready-made model parameters:

  • Top events for the FTA are selected from the project's loss criteria and confirmed fault domains; a catalog row is not automatically a system top event.
  • Basic-event rates may use the evidence in Chapter 14.3 only where component, population, maturity, duty, and event definition match. Otherwise collect project/fleet evidence or preserve the uncertainty.
  • Shared events come from verified topology and operating dependencies. Treat the coupling rows as candidates; do not convert every row into a common-cause term or assign a numeric beta without a justified method and dataset. The model mechanics are in Chapter 12.5.
  • Recovery distributions require measured or justified detection, protection, repair, restart, and validation times. Checkpoint state can couple recovery paths, but it is not universally the dominant term.
  • IST cases come from approved risk-based requirements and use the safest valid evidence method described in Chapter 13.6; the catalog does not authorize a hazardous live event.
Worked example: test whether a coolant leak can become a shared row event

Start with a quick-disconnect leak and trace the actual project topology. Its initial fault domain, detection time, coolant conductivity, electrical exposure, isolation behavior, shared controls, check valves, pump response, and repair access decide whether it remains local or contributes to a branch or row event. Do not assign a one-hour repair or a seconds-to-HBM cascade until those states have been demonstrated or otherwise justified.

Negative-pressure architecture, zoned detection and isolation, qualified leak response, and independent cooling trains can break candidate propagation paths, but each control must be verified for the named configuration. A spoofed dew-point command is a separate adversarial scenario: it may share some protected equipment or consequences, yet has different initiating evidence, detection, persistence, and recovery. Model a shared event only where the confirmed topology actually joins the paths. OT controls are in Chapter 11.10; leak detection and isolation are in Chapter 5.11.

This catalog is the index; the derivations live in their owning chapters. Cooling modes: leak/CDU/pump in Chapter 5.11, controls transients in Chapter 5.12, the density/thermal envelope in Chapter 5.1. Electrical modes: load-step smoothing in Chapter 4.5, grid ride-through in Chapter 4.10, fuel-supply engineering in Chapter 4.9. Connectivity/compute: fiber/latency in Chapter 3.6, optics reliability in Chapter 8.9, the SDC/hard/transient fault taxonomy and AFRs in Chapter 14.3. Water curtailment as a siting gate in Chapter 3.7. Adversarial paths and any validated shared top events are analyzed separately in Chapter 11.10. The catalog feeds the redundancy/fault-domain framework in Chapter 12.1, the goodput-vs-availability rethink in Chapter 12.2, the quantitative RBD/FTA/Monte-Carlo model in Chapter 12.5, and the failure-mode demonstration at Chapter 13.6.
Cite this chapter
Fehn, J. (2026). Failure-Mode / FMEA Catalog (Chapter F). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/appendix-appendices-and-reference-data/f-failure-mode-fmea-catalog (accessed 2026-08-28).
@misc{aidc-F,
  author       = {Fehn, Jacob},
  title        = {Failure-Mode / FMEA Catalog (Chapter F)},
  howpublished = {The Definitive Guide to AI Data Centers},
  year         = {2026},
  url          = {https://aidatacenterguide.com/appendix-appendices-and-reference-data/f-failure-mode-fmea-catalog},
  note         = {Accessed 2026-08-28}
}
Spotted an error? Suggest an edit