2N vs N+1
2N duplicates the whole path; N+1 adds one spare per group. Required maintenance and fault outcomes select between them, and the capex comparison must keep whole-facility, MEP, and duplicated-electrical denominators separate.
| Axis | 2N | N+1 |
|---|---|---|
| Topology | two complete, independent power/cooling paths | one path plus one spare unit per N components |
| Service outcome | no load loss for maintenance or one defined component/path fault; verify common-mode exclusions | maintenance can preserve load, but an unplanned fault outside the covered component set can interrupt service |
| Whole-facility capex | ~25–40% facility planning premium for fault tolerance over concurrent maintainability (single-source heuristic; re-estimate the design) | lower whole-facility cost when fault ride-through is not required |
| MEP-scope capex | Project-specific BOM/TCO premium for 2N over N+1 at the same usable load | the MEP-scope baseline for this comparison |
| Duplicated-electrical scope | project-specific electrical-topology delta; do not infer it from the whole-facility or MEP percentages | one distribution path with spare components avoids a fully duplicated electrical chain |
| Concurrent maintainability | yes — either path carries full load during work | yes only when the component and path arrangement supports maintenance without load loss |
| Failure exposure | single-component faults invisible to IT load | correlated failures or maintenance windows can force load shedding |
| AI-era framing | protects the building, not cluster goodput; measure productive time for the named fleet, job, window, and event accounting | checkpointing, cordons, and spares target measured recovery loss under the same definition |
Quantitative cells are the guide's canonical figures — each is date-stamped and sourced in the numbers register and derived in the chapters below.
How the decision falls
Buy 2N when the required outcome is no load loss during maintenance or one component/path fault and the avoided outage cost exceeds the capex delta; revenue-serving is one consequence when its SLA creates that objective. If checkpoint recovery keeps restart loss within the service objective, goodput investments often outrank a second complete path.
What would flip it: An availability penalty or recovery loss that exceeds the 2N capex delta moves a restart-tolerant service objective toward duplicated paths.
Model this fork with your own numbers: Redundancy & availability calculator →
Full derivations, worked examples, and the numbers behind this matrix: Redundancy topologies and fault domains (Ch 12.1) · Goodput vs facility availability (Ch 12.2) · SLAs and goodput contracts (Ch 12.4)