Appendix F
Failure-Mode / FMEA Catalog
Random failures and adversarial scenarios need distinct initiating evidence and controls; link them only where a validated topology shows that they converge on a shared top event or protective response.
What you'll decide here
- Use this appendix as a lookup, not a narrative: find your failure mode in the master table, then jump to the owning chapter (last column) for the engineering derivation — this catalog summarizes, it does not replace, the canonical treatment.
- Keep random and adversarial initiators, likelihoods, authorities, observability, persistence, controls, and recovery evidence distinct; cross-link only the validated shared top event or protection function.
- Treat the blast-radius column as the input to your fault-domain and RBD work (Chapter 12.1 / 12.5): a mode that strands one rack and a mode that trips the campus POI sit in different availability tiers and deserve different redundancy spend.
- Walk the propagation column for cascade interlocks. Most catastrophic outages here are not the first fault — they are a single fault that defeated a shared mitigation (one CDU, one fuel header, one protection setting) and took an entire fault domain with it.
- Pair detection latency against propagation speed for each mode. Where the fault propagates faster than your detection-plus-actuation loop (load-step→grid trip, HBM runaway), the only viable mitigation is preventive or inertial, never reactive — design accordingly.
This appendix consolidates the failure modes scattered across the engineering chapters into one uniform FMEA register. It is referenced from the resilience-standards chapter (Chapter 12.1), the reliability rethink (Chapter 12.2), the component-failure-rate chapter (Chapter 14.3), and the integrated-systems-test chapter (Chapter 13.6); and it is consumed directly by the quantitative availability model in Chapter 12.5, which draws its top events and common-cause couplings from the rows below. The canonical engineering of each mode lives in the chapter named in the right-most column — this catalog is the index and the cross-walk, not the derivation.
Some random and adversarial scenarios can converge on a similar top event or protective response, but they do not automatically share one record. Keep initiating evidence, likelihood, authority, observability, persistence, propagation, security controls, and recovery distinct. Where the validated topology joins the paths, link the records to the shared top event and reuse the proven protective response; do not claim every random failure has an attacker twin. The cyber-physical analysis is in Chapter 11.10.
How to read a row
Each failure mode is recorded against six fields, applied identically across every table so the catalog is sortable and comparable:
- Trigger — the initiating event (random cause first; the induced/attacker path is noted where it differs materially).
- Propagation path — how the fault spreads, and critically which shared mitigation it defeats to escalate from a local fault to a cascade.
- Detection — the sensing modality and its characteristic latency relative to propagation speed (the decisive ratio).
- Blast radius — the fault domain affected at full propagation: node, rack, row, hall, or campus/POI.
- Mitigation — the preventive or containing control, classified as preventive (stops the trigger), inertial (buys ride-through time), or reactive (acts after detection).
- Recovery — the path back to service and its characteristic time-to-restore (TTR).
Where propagation outruns detection-plus-actuation, the only effective controls are preventive or inertial — which is why that ratio is called out per mode in the notes.
Master FMEA catalog — thermal & mechanical (cooling) modes
| Failure mode | Trigger | Propagation path | Detection | Blast radius | Mitigation | Recovery | Owner |
|---|---|---|---|---|---|---|---|
| Coolant-leak cascade | QD/manifold/cold-plate breach, hose chafe, gasket creep; induced: spoofed dew-point setpoint forcing condensation | Local drip → conductive coolant on busbar/PDU → arc/short → de-rate or trip of the powered branch; if it reaches a shared CDU controller, the whole CDU loop and its row drop together | Floor/leak-rope sensors + CDU flow/pressure-decay; latency seconds-to-minutes; propagation can outrun it on a high-flow breach | Rack → row (if the CDU is shared); hall if isolation valves are absent | Negative-pressure loops (leak draws air in, not coolant out); dripless UQDs; zoned isolation valves; per-rack leak detection; N+1 CDU with independent controllers (preventive + reactive) | Isolate the branch, drain/flush the loop, replace the failed coupling, re-pressure-test, re-fill, re-commission worst-case branch; TTR hours per rack | 5.11 |
| CDU / pump failure | Pump bearing/VFD failure, seal loss, filter blockage, control-board fault; induced: malicious VFD firmware or controller DoS | Flow may be maintained, reduced, redistributed, or lost depending on pump arrangement, check valves, control sequence, loop volume and thermal mass, standby readiness, and IT throttling; temperature response follows that project topology | Pump state, differential pressure, flow, coolant temperatures, valve position and IT telemetry; required detection and trip timing come from the validated transient | The affected hydraulic zone — rack, row, or larger only where the topology shares the failed resource | Qualified duty/standby or N+1 pumps and CDUs, independent power/control where required, validated switchover, adequate thermal/flow ride-through, and coordinated IT throttling | Execute the validated failover or controlled IT response, isolate and repair the failed train, then re-qualify flow and controls; recovery time is topology- and failure-specific | 5.11 |
| Cooling-controls transient excursion | Synchronized GPU load drop (job ends / checkpoint pause) → loop heat input collapses faster than valves/VFDs can slew; setpoint hunt / control-loop oscillation | On a rapid load drop, supply-coolant temp overshoots downward → transient dew-point excursion → condensation risk on cold surfaces; or anti-hunting failure drives sustained oscillation that fatigues actuators and destabilizes neighboring loops | Coolant supply-temp rate-of-change, dew-point margin sensor, valve-position hunting; detectable but the excursion window is brief | Row → hall (controls coupling); condensation risk is local to cold surfaces | Slew-rate limits on control valves and pump VFDs; anti-hunting tuning; dew-point margin floor; coordinate cooling setpoints with the rack BBU/BESS load-smoothing spine (preventive) | Re-tune control loops, restore dew-point margin, dry/inspect any condensation; TTR minutes, no hardware loss if caught | 5.12 |
| HBM thermal runaway | Cold-plate contact loss, TIM pump-out, local flow starvation, or sustained over-temp on a stacked-DRAM site; induced: CDU disablement holding flow at zero | HBM junction temp climbs → ECC error rate rises → uncorrectable error / package damage; on a tightly-coupled training step the failed device stalls the synchronous collective and the whole job stalls behind the straggler | Per-die thermal telemetry, ECC/CE rate trend, GPU throttle flags; trend-detectable early, but runaway is fast once contact is lost | Node (the GPU/HBM package) → job (synchronous training stalls on the straggler) | Thermal screening/burn-in pre-deployment; ECC-rate alarming with proactive drain; flow-failure throttle floor; hot-spare nodes so the scheduler evicts and replaces the straggler (preventive + reactive) | Evict the node, fail the job over to a hot spare, RMA the package; training resumes from last checkpoint (TTR = checkpoint interval + restart) | 14.3 |
A direct-to-chip loop's response to pump or CDU loss is topology-, control-, and inertia-dependent. Pump arrangement, check valves, loop volume and metal mass, accumulator behavior, power ride-through, valve sequence, standby readiness, rack thermal mass, and IT throttling determine whether flow is maintained or lost and how quickly junction temperatures rise. Detection, protective action, and recovery timing must therefore come from the validated project transient, not from a universal seconds-versus-minutes rule. Engineer and test the complete cause-and-effect in Chapter 5.11; feed the measured response into the reliability model in Chapter 12.2.
Master FMEA catalog — electrical & power modes
| Failure mode | Trigger | Propagation path | Detection | Blast radius | Mitigation | Recovery | Owner |
|---|---|---|---|---|---|---|---|
| Simultaneous-GPU-load-step grid trip | Thousands of GPUs ramp in lockstep at job start/stop/checkpoint; di/dt event on every step; induced: malicious power-cap firmware forcing a synchronized step | Aggregate ramp >1,000 MW/s presented to the POI → voltage/frequency disturbance → if the load-smoothing spine is absent, upstream protection or generators see a step they cannot follow → trip | Power-quality metering at the POI, PMU/PQM; fast, but the di/dt event is faster than any reactive control — milliseconds | Campus (POI) → contributes to wide-area grid disturbance | The chip→BBU→BESS smoothing spine (on-package capacitance → rack BBU → facility BESS); software ramp-rate limits and regulated wind-downs; grid-forming inverters (preventive + inertial) | Re-energize per utility ride-through procedure; restore smoothing controls; no hardware loss if the spine held; TTR minutes if ride-through succeeded | 4.5 |
| Utility ride-through / voltage-disturbance event | Grid fault (e.g. 230 kV line fault) causes a voltage sag at the POI; sensitive customer-side protection drops the load | Undervoltage response across facility loads → cumulative ~1,500 MW customer-side load reduction during the July 2024 six-fault NoVA sequence → the reduction itself destabilizes the grid, a self-reinforcing reliability problem NERC flagged at Level 3 | POI relays, undervoltage/under-frequency elements, PMU; the disturbance is sub-cycle to cycles | Campus (full load drop) → wide-area grid | Fault-ride-through settings tuned to stay online through the sag (SEL/relay, UPS, undervoltage-load-retention); reactive/voltage support toward the POI; ride-through posture aligned to the NERC 2026 Level-3 alert's recommended large-load envelope — guidance, not penalty-backed; TPL-001/PRC standards bind transmission planners and generators, not data-center loads (preventive) | Auto-recover as the sag clears if ride-through held; if tripped, sequenced re-energization and load ramp; TTR minutes | 4.10 |
| BESS thermal runaway | Cell defect, overcharge, internal short, or cooling loss in an LFP facility battery; induced: BMS spoof disabling cell balancing/thermal protection | Single cell vents → exothermic chain to adjacent cells → module-level runaway → fire/off-gas if pack-level isolation and venting fail; loss of the BESS also removes the ride-through and load-smoothing it was providing | Cell voltage/temp telemetry, off-gas (H2/CO) detection, BMS fault flags; early-warning gas detection precedes thermal runaway by a useful margin | BESS enclosure → adjacent enclosures/room if propagation isolation fails | LFP chemistry (higher thermal-runaway threshold than NMC); module-level thermal isolation and dedicated venting/deflagration paths; off-gas detection with pre-emptive isolation; physical separation of BESS from IT (preventive + reactive) | Execute the pre-incident response plan: isolate electrically, keep ventilation/deflagration paths clear, monitor off-gas, coordinate fire-service response on the UL 9540A / NFPA 855 basis of design — never open or manually suppress the enclosure ad hoc; replace module/pack; re-commission; TTR hours-to-days; ride-through reverts to UPS/BBU meanwhile | 4.5 |
| Fuel-supply interruption (on-site generation) | Firm pipeline curtailment (correlated cold-snap), valve/compressor failure, or fuel-quality (Wobbe/dew-point) excursion; LNG/CNG storage depletion | Loss of fuel → on-site turbines/engines de-rate or trip → if the site is islanded or grid-import is constrained, generation cannot meet IT load → controlled load-shed or outage | Fuel header pressure, Wobbe-index/dew-point analyzers, tank level, generator load; minutes of warning on slow depletion, immediate on a hard cut | Campus (islanded sites) → partial if grid-import backstops | 'Synthetic-firm' fuel structure (multiple pipelines + interruptible + on-site LNG/CNG storage); dual-fuel switching; fuel conditioning to spec; sized on-site storage for correlated-curtailment duration (preventive) | Switch fuel source / draw down on-site storage, restore generation; coordinate curtailment with curtailable-load agreement; TTR depends on storage sizing vs outage duration | 4.9 |
Master FMEA catalog — connectivity, compute & data-integrity modes
| Failure mode | Trigger | Propagation path | Detection | Blast radius | Mitigation | Recovery | Owner |
|---|---|---|---|---|---|---|---|
| Fiber cut (inter-/intra-DC) | Backhoe/construction strike, conduit failure, or DCI route cut; induced: deliberate cut of an un-diverse route | Loss of a fiber path → if scale-across/DCI is single-routed, the affected campus or training partition is severed → distributed training stalls or a metro site loses connectivity | Optical LOS/LOF alarms, OTDR, BER collapse; immediate at the physical layer | Link → partition/campus (distributed training) or metro site reach (inference) | Physically diverse, geographically separated fiber routes; protected DCI (ZR/ZR+ with restoration); for training, topology that degrades gracefully on a partition loss (preventive + reactive) | Restore over the diverse path automatically; physical splice repair on the cut route (TTR hours-to-days for the splice, seconds for protected failover) | 3.6 |
| Optics flap storm | Marginal transceiver, dirty/over-bent connector, thermal cycling, or firmware bug causing repeated link up/down; induced: thermal attack on the optics environment | One flapping link → routing reconvergence churn / ECMP rehashing → packet loss and tail-latency spikes propagate across the fabric → on a tightly-coupled collective, the flapping link gates the whole all-reduce and tanks MFU | Per-port link-flap counters, BER/FEC-error trend, CRC errors; trend-detectable but a storm builds in seconds-to-minutes | Link → fabric pod (reconvergence churn) → job (collective stalls) | Pre-install optics burn-in / BER screening off the critical path; FEC margin headroom; auto-quarantine of flapping ports; CPO to remove pluggable failure points where deployed (preventive + reactive) | Quarantine and replace the offending transceiver; let routing reconverge; TTR minutes to swap, plus reconvergence | 8.9 |
| SDC corruption event (silent data corruption) | Marginal/defective silicon (a 'mercurial core'), aging, voltage/thermal margin loss producing a wrong result with no error flag; induced: fault-injection on a known-marginal device | A miscomputed value flows silently into gradients/activations → corrupts the model state or an inference result → undetected for hours-to-days, potentially poisoning a checkpoint and forcing a roll-back of all work since the last clean checkpoint | No native hardware flag — requires a dedicated detection program (Fleetscanner periodic, Ripple in-fleet, Hardware Sentinel runtime); latency hours unless instrumented | Node (the mercurial core) → job/model (silent corruption of state) → potentially every downstream consumer of a poisoned checkpoint | Proactive SDC-hunting at scale; redundant/checksummed compute on critical paths; PVF-aware placement; roll-back to a verified-clean checkpoint; quarantine the device (preventive + reactive) | Identify and quarantine the mercurial core, roll back to last verified-clean checkpoint, re-run; TTR = work since clean checkpoint (the reason hyper-frequent checkpointing pays off) | 14.3 |
| Water-curtailment event | Drought-stage ordinance, basin/withdrawal-permit curtailment, or reclaimed-water supply interruption on an evaporative-cooled site (grid-curtailment orders like ERCOT's SB6 regime are an electrical event, not a water one) | Loss/limit of make-up water → evaporative/cooling-tower capacity drops → heat-rejection ceiling falls below IT load → forced de-rate of compute or a curtailment-driven load-shed | Make-up water flow/level, WUE telemetry, basin level, curtailment-order receipt; advance notice on scheduled curtailment, immediate on a hard order | Hall → campus (heat-rejection limited) | Closed-loop / zero-evaporation cooling design (designs water out of the risk); reclaimed/non-potable sourcing; on-site water storage; curtailment-tolerant workload scheduling (batch defers); thermal-storage buffer (preventive) | Shift heat rejection to the closed-loop/dry path or draw down water storage; defer curtailable batch load; TTR = curtailment duration; structural fix is closed-loop conversion | 3.7 |
Three of these four modes share the opposite signature from the cooling and load-step modes: a long detection latency against a slow-burning blast radius. An SDC event can poison a checkpoint and sit undetected for days; an optics flap can quietly erode MFU before anyone correlates the tail-latency to a single marginal transceiver; a fiber cut on a poorly-instrumented diverse path can go unnoticed until the protect path also fails. For this class the decisive investment is detection, not faster actuation: a dedicated SDC-hunting program, per-port flap/FEC telemetry, and verified-clean checkpointing convert an invisible, unbounded-blast-radius fault into a bounded, recoverable one. The empirical fault taxonomy and the AFRs that feed these rows into the availability model are the canonical content of Chapter 14.3; the SDC detection program is treated there and demonstrated at commissioning in Chapter 13.6.
Common-cause couplings & cascade interlocks
For the availability model, the catalog's most useful output is the couplings between modes — the shared resources whose failure makes two independent-looking modes fail together. These are the common-cause terms an RBD or fault tree must capture, or it will badly over-state availability. The table below names the cross-mode interlocks worth modeling explicitly.
| Candidate shared condition | Modes it may couple | Condition to verify | Design / model response |
|---|---|---|---|
| CDU, controller, loop, or power source shared by a hydraulic zone | Leak escalation, loss of flow, thermal protection | Which racks actually share flow, control, isolation and power; whether one fault defeats standby | Model one shared event only for the confirmed fault domain; separate or isolate trains where required |
| One storage system used for smoothing and ride-through | Load-step response, ride-through, storage isolation | Whether the same cells, inverter, controls, protection or bus provide both functions | Separate functions or model the confirmed common dependency and its protection states |
| Fuel source or delivery path exposed to one regional event | Fuel interruption and generation availability | Contract, pipeline, storage and transport independence during the named event duration | Represent the named supply event explicitly; do not infer independence from two contracts |
| Fiber routes sharing conduit, crossing, right-of-way, building entry or carrier equipment | Route cut and loss of protect path | Surveyed physical diversity end to end, including meet-me and power dependencies | Use an explicit route-loss event where sharing is confirmed; correct false diversity |
| Checkpoint policy and last-known-good state | Restart and rollback recovery times | Integrity validation, cadence, write path and corruption horizon | Treat as a recovery-time dependency, not a hardware common-cause beta |
| Shared BMS/DCIM/SCADA authority or communications | Cooling setpoint, flow and load-control excursions | Which commands, sensors, interlocks and overrides share one control plane | Model only confirmed shared authority; provide independent protection where the hazard basis requires |
Using this catalog in the availability model
The rows above are inputs to reliability work, not ready-made model parameters:
- Top events for the FTA are selected from the project's loss criteria and confirmed fault domains; a catalog row is not automatically a system top event.
- Basic-event rates may use the evidence in Chapter 14.3 only where component, population, maturity, duty, and event definition match. Otherwise collect project/fleet evidence or preserve the uncertainty.
- Shared events come from verified topology and operating dependencies. Treat the coupling rows as candidates; do not convert every row into a common-cause term or assign a numeric beta without a justified method and dataset. The model mechanics are in Chapter 12.5.
- Recovery distributions require measured or justified detection, protection, repair, restart, and validation times. Checkpoint state can couple recovery paths, but it is not universally the dominant term.
- IST cases come from approved risk-based requirements and use the safest valid evidence method described in Chapter 13.6; the catalog does not authorize a hazardous live event.
Worked example: test whether a coolant leak can become a shared row event
Start with a quick-disconnect leak and trace the actual project topology. Its initial fault domain, detection time, coolant conductivity, electrical exposure, isolation behavior, shared controls, check valves, pump response, and repair access decide whether it remains local or contributes to a branch or row event. Do not assign a one-hour repair or a seconds-to-HBM cascade until those states have been demonstrated or otherwise justified.
Negative-pressure architecture, zoned detection and isolation, qualified leak response, and independent cooling trains can break candidate propagation paths, but each control must be verified for the named configuration. A spoofed dew-point command is a separate adversarial scenario: it may share some protected equipment or consequences, yet has different initiating evidence, detection, persistence, and recovery. Model a shared event only where the confirmed topology actually joins the paths. OT controls are in Chapter 11.10; leak detection and isolation are in Chapter 5.11.
Cite this chapter
Fehn, J. (2026). Failure-Mode / FMEA Catalog (Chapter F). The Definitive Guide to AI Data Centers. https://aidatacenterguide.com/appendix-appendices-and-reference-data/f-failure-mode-fmea-catalog (accessed 2026-08-28).
@misc{aidc-F,
author = {Fehn, Jacob},
title = {Failure-Mode / FMEA Catalog (Chapter F)},
howpublished = {The Definitive Guide to AI Data Centers},
year = {2026},
url = {https://aidatacenterguide.com/appendix-appendices-and-reference-data/f-failure-mode-fmea-catalog},
note = {Accessed 2026-08-28}
}