The Definitive Guide toAI Data Centers
Ask the GuideAboutAccount
CalculatorsInference $/M tokens

Inference $/M tokens

Compare rented and owned inference capacity in $/M tokens on the same throughput and productive-use assumptions, with the price-decline context you should underwrite against.

ScenarioNVIDIA GB200 NVL72 · 56 active / 72 installed GPUs · 0.13 MW IT · $5M programguide defaults (asof 2026-07)Edit the shared scenario →

Inference cost per million tokens

$0.82 / M tokens owned · $5.32 rented
Owned capacity at 70% productive utilization$0.82/M tokens
Rented capacity at the same 70% productive utilization$5.32/M tokens

Compare ownership and rental on the same throughput and productive-use denominator. Owned node cost is the calendar-hour allocation of capex, energy, and opex; rented node cost is the capacity held for that hour. Market self-serve fell ~$10 → ~$2.50 / M tokens in a year (~4×), so underwrite inference with a price-decline curve → Ch 1.8.

These are transparent screening estimates, not final designs, financial advice, or project approvals. Replace every default with current project data and have the responsible project authorities validate the result. You can save, share a permalink, or export to CSV; see the full calculator suite.