The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount

Calculators

Open models for the numbers that decide an AI build — one project, worked end to end: size the cluster, price the build, cost the runtime, then see if the finance closes. Defaults trace to the numbers register (current as of 2026-07); change them to your case, then save, share a permalink, or export to CSV.

Shared starting assumptions
Connected fields: Cluster sizing — model → MW (17/17) · Build cost — $/MW capex (4/5) · Training run — cost & time (12/12) · Project finance — CFADS, DSCR & IRR (17/17) · GPU TCO & $/GPU-hour (11/11) · Inference $/M tokens (4/4) · Facility energy & water (4/4) · Rack cooling feasibility (2/7) · Reliability requirements (2/2). Independent: Site-scoring playbook (0/6; local judgment inputs).
Workload
Commercial boundary: the merchant factor drives rental revenue; productive utilization drives owned/rented unit-cost comparisons. They are parallel analytical cases, not additive hour buckets or a hybrid allocation.
Hardware generation
1200 W · 186 GB HBM · 72/rack · 132 kW/rack · 2.5 PFLOPS dense · 72-GPU indivisible SKU minimum
Performance profile: llama-3.3-70b · fp8 · 32,000 context · 64 batch/concurrency · guide assumption (11,200 tok/s); substitute a measured value from this exact profile
Facility & location
Financing & schedule
NVIDIA GB200 NVL7256 active / 72 installed GPUs · 16 stranded132 kW/rack → Named GB200 NVL72 direct-liquid configuration0.13 MW IT / 0.15 MW facility$5M program$4.14/active GPU-hr · $0.82/M tok owned at 70% productive utilization511 t CO₂/yr64.9% equity IRR · 3.06× min DSCR
All assumptions are the guide defaults (asof 2026-07). Everything runs in your browser — the share link carries the scenario in the URL fragment (never sent to a server); only an explicit Save stores it to your account, versioned (v4).

Size it

Cluster sizing — model to megawatts

56 active / 72 installed GPUs · 0.15 MW design
VRAM per replica (weights 70 + KV 676 GB)858 GB
Qualified GPUs per replica → replicas for demand (VRAM floor 5)8 × 7
Installed GPUs / stranded slots72 / 16
Installed racks (72/rack)1
Facility design draw on installed units (incl. host + PUE)0.15 MW
Rental cost on active GPUs$767,318/mo

VRAM establishes only a memory floor. Enter a replica degree qualified for the named model, runtime, hardware, context and batch profile; supported degrees need not be powers of two. Purchase quantum is a separate SKU or contract input and is not inferred from rack density. Facility design power follows installed units, while rental cost follows the active fractional fleet. At long context the KV cache can dominate VRAM → Ch 9.7.

Rack cooling feasibility

direct liquid cooling with air removal of the declared residual heat is feasible
Current liquid / residual-air split114.8 / 17.2 kW
Future residual-air load17.2 kW vs 20 kW limit

This is a preliminary envelope screen, not equipment qualification. Rack kW is one input alongside component heat flux, the named OEM heat split, airflow/inlet limits, supported TCS/FWS or entering-water conditions, room residual rejection, climate, serviceability, redundancy, and the refresh tail. Use product capacity only at its stated air/water temperatures, flow, pressure, fan state, containment, and heat-capture target → Ch 5.4.

Build it

Build cost — capex per megawatt

$2M excl. servers
Facility (shell, electrical, cooling, land)$2M
Network & cluster infrastructure$1M
Intensity, servers excluded$16.7/W

Benchmarks: JLL global-average shell/core ≈ $11.3M/MW (50 MW single-tenant, air-cooled; liquid +10%); Epoch's 1 GW model runs facility+land+utility ~$11.8M/MW and network ~$4.9M/MW, with servers ~$21.2M/MW on top → ~$38B/GW up-front. Electrical dominates the facility split → Ch 2.5, Ch 1.8.

Scope & how the benchmarks reconcile

This models construction-period capex excluding the server fleet. The benchmarks, reconciled on Epoch AI's May-2026 1 GW model: $11.3M/MW is JLL's global-average shell-and-core (air-cooled, 50 MW single-tenant; liquid cooling adds ~10%); Epoch's facility + land + utility works line is ~$11.8M/MW and its network and cluster infrastructure ~$4.9M/MW; servers — GPUs included — add ~$21.2M/MW, which is how a 1 GW campus reaches ~$38B up-front. Electrical systems are 45–70% of construction cost (as of 2026 → register). The server fleet is deliberately outside this calculator — the TCO model owns it, and adding it here would double-count.

Site-scoring playbook

64 / 100 · Workable with mitigation

Kill gates come first: power or interconnect at ≤2, or water or land/zoning at ≤1, disqualifies a site no matter how the weighted score averages out. Weights reflect the reordered 2026 hierarchy — power availability now leads, ahead of latency and land. Tune to your workload (training tolerates latency; inference doesn't) → Ch 3.13.

Reliability topology requirements

Topology not determined
Concurrent maintenanceRequired
Single-component fault ride-throughNot required by this screening input

Project topology requires named maintenance and fault states, the exact affected path, transfer interruption, post-event loading, path and control independence, common-mode dependencies, a recovery SLO, and workload and contract consequences. These two choices are useful requirements, but they cannot select N, N+1, distributed-redundant, 2N, a Tier, or a capex premium. Compare and price candidate designs only after that project evidence closes → Ch 12.1.

Run it

Training run — time, cost, energy & CO₂

8.1 days · $36.5M rental
Compute (6·P·T)6.3×10²⁴ FLOPs
GPU-hours1.94M
Facility energy (incl. host + PUE)4,100 MWh (~$0.33M)
Emissions at 384 gCO₂/kWh1,574 t CO₂

MFU and goodput compound: a 40% MFU × 90% goodput scenario yields 36% of peak theoretical FLOPs as committed training work under the declared definitions. Treat 90% versus 96% only as an illustrative sensitivity; measure the named fleet, job, window, and event accounting, then attribute badput before assigning any gap to checkpointing or cordon policy → Ch 12.2.

GPU TCO & cost-per-GPU-hour

$4.14 / active GPU-hr
Annual cost / active GPU$25,357
· Depreciation$18,917
· Energy$1,900
· Opex$4,540
· Installed-to-active allocation1.285714× across all three cost lines
Breakeven utilization vs rental15%

Ownership beats the rental rate only above the breakeven utilization computed from your own inputs — below it a debt-financed cluster bleeds cash → Ch 1.8.

Inference cost per million tokens

$0.82 / M tokens owned · $5.32 rented
Owned capacity at 70% productive utilization$0.82/M tokens
Rented capacity at the same 70% productive utilization$5.32/M tokens

Compare ownership and rental on the same throughput and productive-use denominator. Owned node cost is the calendar-hour allocation of capex, energy, and opex; rented node cost is the capacity held for that hour. Market self-serve fell ~$10 → ~$2.50 / M tokens in a year (~4×), so underwrite inference with a price-decline curve → Ch 1.8.

Facility energy & water

Facility draw0.15 MW
Annual energy1,330 MWh
Annual energy cost$106,381
Annual water0.5 M L

PUE bands: legacy air 1.4–1.6 · direct-to-chip liquid 1.05–1.15 → Ch 15.1.

Fund it

Project finance — CFADS, DSCR, IRR

64.9% equity IRR
Project (unlevered) IRR44.5%
Min / avg DSCR (CFADS ÷ debt service)3.06× / 4.55×
Debt at COD (incl. $0M IDC)$3M
Equity invested (incl. $0M DSRA)$2M
Project NPV @ 10%$14M
Equity multiple13.78×

A real CFADS pro-forma: capex draws over the build with capitalized IDC, a revenue ramp, cash taxes net of the straight-line depreciation and interest shields, and sustaining capex. DSCR is CFADS ÷ debt service — the lender's ratio, stricter than EBITDA coverage; screen against ~1.3–1.5× contracted, ~1.75–2× merchant. Contracted offtake supports more leverage than merchant. Working-capital swings and NOL carryforwards are not modeled → Ch 2.5.

Then turn the sizing into dates: the lead-time planner reverse-schedules every long-lead PO from your ready-for-service target.