Glossary
618 terms across power, cooling, compute, networking, reliability, and economics — the vocabulary this field borrowed from six industries, defined the way the guide actually uses it. Term pages pull their current numbers live from the register.
618 of 618 terms
$/GPU-hr
The all-in cost to operate one accelerator for one hour; the core unit economic for build-versus-rent decisions.
$/M-tokens
The cost to serve one million tokens, the revenue-side unit economic for inference businesses.
24/7 CFE · Carbon-Free Energy
Matching every hour of consumption with carbon-free generation, a stricter goal than annual renewable matching.
2N
A capacity/path arrangement with two independent systems, each sized to carry the declared design load. 2N describes topology and capacity, not zero impact or a guaranteed availability outcome; shared dependencies and operating state still matter.
800 VDC
Direct-current rack distribution for megawatt-class racks that cuts conversion stages and copper versus traditional AC.
AALC · Air-assisted liquid cooling
A named air-assisted liquid-cooling system uses an internal liquid loop and air-side heat rejection without facility water at the rack. Capacity and suitability depend on the named equipment, rack heat split, room airflow/inlet conditions, and facility heat-rejection envelope—not a universal air-cooling ceiling.
ABS · Asset-Backed Securitization
Financing that bundles cash-flowing assets (such as leases or GPUs) into securities sold to investors.
AEC · Active Electrical Cable
Copper cable with a redriver or retimer in the connector, buying a few metres of reach past passive DAC at a power and reliability cost; it sets the distance where optics become unavoidable.
Aeroderivative turbine
A jet-engine-derived gas turbine that starts fast and ramps quickly, suited to on-site or backup power.
AFR · Annualized Failure Rate
The fraction of a component population expected to fail per year; the fleet-planning number behind spares pools and maintenance budgets.
AHJ · Authority Having Jurisdiction
The local official who interprets and enforces adopted codes; their reading, not the code text, decides whether a novel cooling or battery design gets signed off.
Air-gap
Physically isolating a system or network from any external connection to protect highly sensitive workloads.
All-gather
A collective that assembles each GPU's data shard onto every GPU, common in sharded training and inference.
All-reduce
A collective operation that sums gradients across all GPUs and shares the result, the dominant traffic in training.
ANSI · American National Standards Institute
The US standards accreditor whose imprimatur turns industry documents into national standards that lenders, insurers, and code officials treat as the governing reference.
Anycast
Routing that sends a request to the nearest of several sites sharing one address, used for low-latency global serving.
Approach temperature
The gap between a coolant's temperature and the medium it rejects heat to; a small approach demands a bigger exchanger.
Archetype
A reference workload pattern (pretraining, post-training, RL, online or batch inference, edge) that drives design choices.
Arithmetic intensity
The ratio of compute operations to bytes moved; it determines whether a kernel is compute- or memory-bound.
ASCE · American Society of Civil Engineers
The body behind the loading standard that assigns a facility its Risk Category and sets seismic anchorage and flood-elevation requirements — structural costs fixed at siting.
ASHRAE
The engineering society whose thermal guidelines define the temperature and humidity envelopes for IT equipment.
ASHRAE TC 9.9
The industry committee whose A1-A4 air and W17-W45 water classes define the temperature envelopes IT gear can run in.
ASIC · Application-Specific Integrated Circuit
A chip hard-wired for one task; AI ASICs trade flexibility for efficiency versus general-purpose GPUs.
ASME · American Society of Mechanical Engineers
The body behind the B31 pressure-piping code family that governs how North American liquid-cooling loops are designed, welded, inspected, and pressure-tested.
ASTM · ASTM International
The standards body whose environmental site assessment practice defines the Phase I ESA a buyer must follow to earn liability protection on a land purchase.
ATS · Automatic Transfer Switch
A switch that automatically moves load from utility to backup generator when grid power fails.
Autoregressive
Generating output one token at a time, each conditioned on all previous tokens, the basis of LLM decoding.
Availability zone · AZ
An isolated data-center location within a cloud region, designed to fail independently of its peers.
AVR · Automatic voltage regulator
The generator control that holds terminal voltage through a load step; if its response does not match the stability-study model, the ride-through case is wrong.
B200
NVIDIA's Blackwell-generation GPU as shipped in HGX B200 air- or liquid-cooled nodes — the tier below the flagship NVL72 rack systems.
B300
NVIDIA's later Blackwell-generation GPU, the memory-expanded step above B200 and a reference point for how per-GPU HBM capacity and rack power are climbing each generation.
BACnet · Building Automation and Control Network
The building-automation protocol carrying setpoints and alarms between the BMS and plant; unauthenticated by design, so segmentation rather than the protocol provides its security.
Badput
Compute spent on work that is wasted (failed, redundant or thrown away), the opposite of goodput.
BBU · Battery Backup Unit
Rack-level battery that rides through power dips and absorbs sudden GPU load swings before the BESS or genset reacts.
Behind-the-meter · BTM
On-site generation connected on the customer side of the utility meter, bypassing grid interconnection delays.
BER · Bit-Error Rate
The fraction of transmitted bits received in error; a key quality metric for high-speed optical and copper links.
BES · Bulk Electric System
The transmission-level system whose registered entities carry NERC reliability and CIP obligations; a campus large enough to clear the registry criteria inherits that compliance burden.
BESS · Battery Energy Storage System
Facility-scale battery bank that smooths load transients, provides ride-through, and can bridge to generators.
BF16 · Brain Float 16
A 16-bit floating-point format with a wide exponent range, a common default for stable AI training.
BGP · Border Gateway Protocol
The routing protocol of both the internet and modern leaf-spine fabrics; distributes reachability between switches and to the outside world.
BICSI · Building Industry Consulting Service International
The telecom-infrastructure association whose data-center design and operations standard offers an alternative resilience-class framework to the Uptime Tier ladder.
Bisection bandwidth
The aggregate bandwidth across the worst-case cut of a network; the key figure for all-reduce-heavy training.
Black start
Restarting generation and energizing a grid or site from a complete shutdown without external power.
Black-building test · Pull-the-plug test
A commissioning test that cuts utility power to prove backup generation and transfer carry the full load.
Blast radius · Fault domain
The scope of impact when something fails; good design shrinks the blast radius so one fault affects few resources.
Blowdown
Periodically draining concentrated mineral-laden water from a cooling tower to control scale, a key water loss.
BlueField
NVIDIA's DPU line, terminating tenant networking, encryption, microsegmentation, and storage initiation in hardware at the host edge — an enforcement point bought at server-spec time or never had.
BMC · Baseboard Management Controller
An always-on chip that remotely monitors and manages a server's hardware independent of its CPU and OS.
BMS · Building Management System
The control system that runs a facility's mechanical and electrical equipment such as cooling, power and alarms.
BOD · Basis of Design
The design engineer's documented account of how the proposed systems satisfy each owner requirement; the commissioning authority verifies it answers the OPR before construction, not after.
BOM · Bill of Materials
The itemized list of every part in a system or build; design changes ripple through it into cost, lead times, and power budgets.
Breakeven utilization
The fraction of capacity that must be sold or used for a facility's revenue to cover its costs.
Bring-up
The process of powering, validating and tuning a new cluster until it passes acceptance and can run real jobs.
Brownfield
A project that reuses or retrofits an existing building or site rather than building new.
Build-to-suit · BTS
A data center custom-built and leased to a specific tenant's specifications, usually under a long-term contract.
Burn-in
Running new hardware hard for a period to surface early ('infant mortality') failures before production use.
Busbar
A solid metal conductor distributing high current within a rack or system, replacing bulky cabling.
Busway
Overhead enclosed busbar run with tap-off boxes that distributes power flexibly across rows of racks.
BYOP · Bring Your Own Power
Self-supplying generation when the grid cannot commit to an energization date, trading interconnection wait for fuel supply, emissions permitting, and stranded-asset risk.
CAB · Change Advisory Board
The board authorizing higher-risk changes; on an AI campus it has to span facility and cluster, because a change to either can take down the other.
CAGR · Compound Annual Growth Rate
The smoothed yearly growth rate of a quantity over a period, used to project demand, cost or capacity.
Caliptra
An open-source silicon root-of-trust block letting chips verify their own firmware, backed by OCP and hyperscalers.
Capex · Capital Expenditure
Up-front spending on long-lived assets like buildings, power gear and GPUs, depreciated over their useful life.
Cascade-to-inference
Repurposing older training GPUs for inference as newer chips take over training, extending hardware economic life.
CBA · Community Benefits Agreement
A binding contract funding host-community priorities as the price of approval; distinct from a tax abatement, which reduces host revenue rather than adding to it.
CCGT · Combined-Cycle Gas Turbine
A high-efficiency gas power plant that reuses turbine exhaust heat to drive a steam turbine, a behind-the-meter option.
CDN · Content Delivery Network
A distributed network of caching servers that delivers content from locations close to users.
CDU · Coolant Distribution Unit
Heat exchanger plus pumps isolating the clean technology loop from facility water and providing leak containment.
CFADS · Cash Flow Available for Debt Service
The project cash flow lenders size debt against and the numerator of the coverage ratio; the central output of a campus project-finance model.
CFD · Computational Fluid Dynamics
Simulation of airflow and heat that engineers use to design and verify data-center cooling.
Chain-of-thought · CoT
Prompting or training a model to reason in explicit intermediate steps, trading more tokens for better answers.
Checkpoint
A periodic save of model weights and optimizer state so a long training run can resume after a failure.
Chiplet
A smaller die combined with others in one package, letting designers mix processes and beat single-die size limits.
Chunked prefill
Breaking a long prompt's prefill into chunks interleaved with decode so latency stays steady under load.
Circular financing
Arrangements where a chip vendor invests in customers who use the money to buy its chips, raising scrutiny over demand.
ClusterMAX
SemiAnalysis's rating system grading GPU cloud providers on reliability, performance and operational maturity.
CMBS · Commercial Mortgage-Backed Securities
Securitized commercial real-estate debt, one of the two capital-markets channels through which stabilized data-center assets get refinanced at scale.
CMMC · Cybersecurity Maturity Model Certification
The US Defense Department framework certifying contractors' cybersecurity to handle controlled unclassified information.
CMMS · Computerized maintenance management system
The asset and preventive-maintenance system of record; having it loaded with assets and schedules is an operational-readiness gate, not a post-go-live cleanup task.
CNG · Compressed natural gas
Natural gas stored on site under pressure; with LNG and dual-fuel switching it is how an interruptible pipeline supply is synthesized into firm fuel for on-site generation.
CNP · Congestion Notification Packet
The RoCE feedback packet a receiving NIC returns after seeing an ECN mark; a rising CNP rate is the earliest sign congestion control is throttling, ahead of any PFC pause.
COD · Commercial Operation Date
The date a facility or power asset begins commercial service, often a contractual and financing milestone.
Cold plate
A liquid-cooled metal block clamped to a chip that conducts heat into the coolant in a direct-liquid-cooling loop.
Collective
A coordinated communication pattern across many GPUs (all-reduce, all-gather, reduce-scatter) central to distributed AI.
Colocation · Colo
Renting space, power and cooling in a shared facility, in wholesale (large blocks) or retail (rack-level) form.
Commissioning · Cx
The structured testing (levels L1-L5) that proves a facility works correctly under load before it goes live.
Concurrent maintainability
The ability to maintain or replace any component without shutting down IT load, the defining trait of Tier III.
Confidential computing
Protecting data while it is being processed by running it inside a hardware trusted execution environment.
ConnectX
NVIDIA's network adapter line, carrying the RDMA, hardware timestamping, and reordering features a fabric's transport depends on; its firmware is a first-class item in cluster version lockstep.
Containment
Physically separating hot and cold air (cold-aisle, hot-aisle, or chimney) so cooling air is not wasted by mixing.
Continuous batching · In-flight batching
Dynamically adding and removing requests from a running inference batch to keep the GPU busy and lift throughput.
Cooling cliff · Density wall
The rack-power point (~100 kW) above which air cooling fails and liquid cooling becomes mandatory.
Cooling distribution · L2L / L2A
Heat-exchange schemes moving heat liquid-to-liquid (L2L) or liquid-to-air (L2A) between cooling loops.
Cooling tower
A structure that rejects heat by evaporating water; efficient but the main driver of data-center water consumption.
COP · Coefficient of Performance
A heat pump or chiller's ratio of heat moved to electricity consumed; higher means more efficient cooling or heating.
Cordon and drain
Marking a node unschedulable and moving its work off so it can be serviced without disrupting the cluster.
CoreWeave
A GPU-specialist cloud provider whose published cluster performance and contract structures are common reference points for AI-infrastructure economics.
CoWoS · Chip-on-Wafer-on-Substrate
TSMC's 2.5D packaging that joins logic die and HBM stacks on a silicon interposer; its wafer capacity gates AI supply.
CPO · Co-Packaged Optics
Optics integrated into the switch or accelerator package to beat copper-reach limits at the cost of serviceability.
CPU · Central Processing Unit
The host processor in a GPU node: runs the OS, storage and data pipelines, and orchestration while the accelerators do the tensor math; the CPU:GPU ratio is a node-design choice.
CRAC · Computer Room Air Conditioner
A refrigerant-based room cooling unit; the legacy air-cooling workhorse now giving way to liquid for AI density.
CRAH · Computer Room Air Handler
A chilled-water room air handler that cools the data hall, more efficient than refrigerant-based CRAC units.
Critical path · CPM
The longest chain of dependent tasks whose any slip delays the whole project; everything off it has float.
CRR · Congestion Revenue Right
A market instrument hedging the price difference between your delivery node and the trading hub — the standard tool against nodal basis risk.
CSRD · Corporate Sustainability Reporting Directive
The EU regime requiring audited sustainability disclosure, which turns efficiency, carbon, and water figures from marketing claims into statements an auditor will test.
CUDA
NVIDIA's programming platform for general-purpose GPU computing; the software moat underpinning its AI dominance.
CUE · Carbon Usage Effectiveness
Kilograms of CO2-equivalent emitted per kWh of IT energy; the carbon companion to PUE.
CUI · Controlled Unclassified Information
A US government data category that is unclassified but protected; handling it pulls in US-persons access control, separated enclaves, and third-party certification.
Curtailable load
Load a site agrees to reduce on the grid operator's signal, trading interruptions for faster or cheaper interconnection.
Curtailment
Forced reduction of a load's or generator's output, often to manage grid constraints; AI loads may trade it for speed.
CVE · Common Vulnerabilities and Exposures
The public identifier for a disclosed vulnerability; GPU-stack, container-runtime, and BMC entries drive the patch cadence a fleet has to sustain without killing jobs.
CVSS · Common Vulnerability Scoring System
The standard severity score attached to a disclosed vulnerability, used to decide which flaws justify an out-of-cycle firmware or runtime patch.
CxA · Commissioning Agent
The independent party that plans and verifies commissioning to confirm the facility performs as designed.
CXL · Compute Express Link
A cache-coherent interconnect over PCIe that lets CPUs, accelerators and memory pools share memory across devices.
D2C · Direct-to-chip
Trade shorthand for cold-plate cooling that brings coolant to the package; the default for frontier-density racks, with two-phase variants as the open successor question.
DAC · Direct Attach Copper
A copper cable carrying high-speed signals over short reach, cheaper than optics for in-rack and adjacent links.
DALI · Data Loading Library
NVIDIA's GPU-accelerated data-loading and preprocessing library, used with sharded formats and GPUDirect Storage to keep accelerators fed rather than stalled on input pipeline work.
DAOS · Distributed Asynchronous Object Storage
An open-source high-performance storage system built for NVMe and persistent memory in HPC and AI clusters.
Data gravity
The tendency for large datasets to attract compute and services, making data expensive and slow to move.
Data parallelism · DP
Replicating the model across GPUs that each process different data and synchronize gradients via all-reduce.
Data residency
The requirement that data be stored and processed within a specific country or jurisdiction.
Data sovereignty
The principle that data is subject to the laws of the nation where it is collected or stored.
DC-DC · Direct-current to direct-current converter
The conversion stage stepping a DC bus down to board rails; collapsing the chain onto a single high-ratio DC-DC stage is the efficiency argument for facility-level DC.
DC-SCM · Datacenter Secure Control Module
The modular card carrying the BMC, root of trust, and TPM off the motherboard, standardizing the management and attestation subsystem across server generations.
DCD · Data Center Dynamics
A data-center industry trade publication, cited for equipment lead times, deployment announcements, and market coverage.
DCGM · Data Center GPU Manager
NVIDIA's tool for monitoring, diagnosing and health-checking GPUs across a fleet.
DCI · Data Center Interconnect
The long-haul links and equipment connecting separate data-center sites, increasingly used to scale AI across buildings.
DCIM · Data Center Infrastructure Management
Software that monitors and manages a facility's power, cooling, space and assets in one place.
DCQCN
The congestion-control algorithm tuning RoCE flows using ECN and PFC; mis-tuned it causes victim flows and stalls.
DDN · DataDirect Networks
A storage vendor whose Lustre-based appliances are a standard certified hot tier in GPU-vendor reference architectures.
DDR5 · double data rate 5 (SDRAM)
The current socketed commodity DRAM generation used for host memory; it competes with HBM for the same DRAM wafer supply, so HBM demand raises its price.
DDTL · Delayed-Draw Term Loan
A loan committed up front but drawn in stages as construction milestones are hit, matching financing to capital needs.
Decode
The memory-bandwidth-bound phase that generates output tokens one at a time after prefill.
DeepSpeed
A distributed training library whose ZeRO optimizer shards model states across devices; standardizing on it determines which parallelism layouts your stack can express without rewriting models.
Delta-T
The temperature rise of coolant across a cold plate or heat exchanger; it sizes flow rate and the warm-water loop.
Demand response
Reducing or shifting electricity use on the grid operator's request in exchange for payments or cheaper rates.
Depreciation
Spreading an asset's cost over its useful life; book life and economic life can differ and reshape reported margins.
Design-basis
The frozen set of design assumptions and requirements that, once signed, lets long-lead gear be ordered.
DGX
NVIDIA's fully integrated AI server and SuperPOD reference system sold as a turnkey appliance.
DH · District heating
A municipal hot-water network able to absorb data-center waste heat; the pre-existing pipe in the ground, not the data center, is the asset that makes reuse bankable.
DICE · Device Identifier Composition Engine
A layered hardware identity scheme deriving each boot stage's key from the one before, so attestation reveals what firmware actually ran on that silicon.
Dielectric fluid
A non-conductive liquid (such as engineered fluorocarbons) used in immersion cooling so it can contact electronics safely.
Digital twin
A live software model of a physical facility or system, used to validate designs and optimize operations.
DIMM · dual in-line memory module
The socketed module that carries host DRAM; in a liquid-cooled AI rack DIMMs are part of the residual air-cooled load and a prime harvest target at decommission.
Disaggregated serving · P/D disaggregation
Running prefill and decode on separate GPU pools so each scales independently for better efficiency.
Distributed-redundant · 3N/2
A redundancy scheme (e.g. 3N/2, 4N/3) spreading reserve capacity across multiple paths for efficiency over pure 2N.
DLC · Direct Liquid Cooling
Direct-to-chip liquid cooling: cold plates remove a named share of component heat through an approved coolant loop. Selection depends on the equipment heat split, component heat flux, airflow/inlet limits, and facility water/rejection envelope—not a universal rack-kW cliff or default.
DMTF · Distributed Management Task Force
The industry body defining Redfish, the schema-based management interface replacing IPMI and now extending into liquid-cooling and facility telemetry.
DMZ · Demilitarized Zone
The semi-trusted zone between untrusted networks and internal systems; on an AI campus it is where brokered vendor access and outbound telemetry belong, segmented hardest away from weights.
DNS · Domain Name System
The name-resolution layer; because regional failover and traffic draining depend on it, it is a shared global dependency capable of turning one region's fault into an estate-wide outage.
DoS · Denial of Service
Starving others of a resource, deliberately or not; on shared GPU nodes a noisy or hostile tenant can deny service without ever breaching isolation.
DPU · Data Processing Unit
A programmable NIC (e.g. BlueField) that offloads networking, storage and security from the host CPU.
DR · Disaster Recovery
The plan and infrastructure for surviving the loss of a site or region: replication, failover, and the recovery-time and recovery-point targets contracted to tenants.
DRA · Dynamic Resource Allocation
The Kubernetes mechanism that replaces counting opaque devices with claiming them by attribute, letting a workload request GPUs by memory, topology, or partition rather than by integer count.
DRAM · Dynamic Random-Access Memory
Volatile system memory attached to the CPU; the capacity tier behind the GPU's on-package HBM, used for caches, KV offload, and data staging.
Dry cooler
A finned coil that rejects heat to ambient air without evaporating water, saving water at the cost of efficiency on hot days.
DSCR · Debt-Service Coverage Ratio
Operating cash flow divided by debt payments; lenders require it above a threshold to ensure loans are serviceable.
DWDM · Dense Wavelength Division Multiplexing
Multiplexing many wavelengths onto one fiber pair; it is what turns leased or owned dark fiber into the aggregate inter-site bandwidth a multi-campus training run needs.
DWPD · Drive Writes Per Day
How many times an SSD's full capacity can be overwritten daily over its warranty, a key endurance rating.
E1.S
A compact NVMe SSD form factor for dense storage and compute nodes; the choice between it and the larger E3.S fixes per-drive bandwidth, drive count, and shelf cooling.
EBITDA
Earnings before interest, taxes, depreciation and amortization; a proxy for operating cash generation.
ECC · Error-Correcting Code
Memory protection that detects and corrects bit errors; uncorrectable ECC errors signal failing memory.
ECMP · Equal-Cost Multi-Path
Standard Ethernet load balancing that hashes each flow onto one of several equal-cost links; hash collisions on a few elephant flows are why AI fabrics add packet spraying or adaptive routing.
ECN · Explicit Congestion Notification
A signaling method that marks packets to throttle senders before congestion forces drops, key to lossless AI Ethernet.
Edge inference
Running AI models close to where data is generated, at the network edge, to reduce latency and backhaul.
EDP · Energy-Delay Product
Energy multiplied by latency per operation; a chip figure of merit penalizing designs that are slow or power-hungry.
EDPp · Electrical Design Power, peak
The peak electrical power a chip or rack can transiently draw — sizes the upstream electrical chain, vs TDP's steady-state thermal envelope.
EED · Energy Efficiency Directive
The EU Energy Efficiency Directive requires qualifying data centers to submit specified energy and sustainability data to the European database. Individual submissions are confidential; the Commission publishes aggregated Member-State/Union indicators, not a public database of facility records.
Egress
Data leaving a cloud or region, typically billed at a premium and a major hidden cost in AI pipelines.
Elastic training
Training that can continue at a reduced GPU count when nodes fail and absorb them back when restored.
Embodied carbon
The greenhouse-gas emissions from making and building hardware and facilities, separate from operating energy.
EMC · Electromagnetic compatibility
The discipline of keeping equipment from emitting or absorbing enough electromagnetic energy to disturb its neighbors; it governs bonding, segregation, and cable routing around high-speed links.
EN 50600
The European data-center facility standard, with Availability Classes paralleling the Uptime Tier scheme.
EnEfG · Energieeffizienzgesetz (German Energy Efficiency Act)
Germany's efficiency law, which converts heat reuse from optional economics into a minimum obligation for new data centers above a power threshold.
EOP · Emergency Operating Procedure
The pre-written response to a failure, executed under time pressure; if it is not written and rehearsed before go-live, the first real event becomes improvisation.
EPC · Engineering, Procurement and Construction
A delivery model where one contractor designs, buys and builds the project, often under a fixed price.
EPMS · Electrical Power Monitoring System
A system that continuously monitors and records the facility's electrical distribution for reliability and analysis.
EPO · Emergency power off
The hardwired kill of facility power; its interlocks are both a life-safety requirement and one of the highest-consequence mis-operation risks in a live hall.
EPRI · Electric Power Research Institute
The utility-funded research institute whose flexibility and generation-cost work shapes how utilities study, price, and queue large data-center loads.
Erasure coding
Splitting data into fragments plus parity so it survives multiple drive failures using far less overhead than full copies.
ERCOT
The grid operator for most of Texas, a frequent AI-data-center destination known for fast interconnection and volatility.
ERF · Energy Reuse Factor
Fraction of facility energy exported as useful heat (e.g. district heating); a rare metric where higher is better.
ESA · Environmental Site Assessment
The Phase I/II investigation establishing a site's contamination status; it gates acquisition diligence at the front end and the environmental closeout that ends decommissioning liability.
ETTR · Effective Training Time Ratio
The share of wall-clock time a training job spends making forward progress rather than restarting, checkpointing, or waiting on failures; the honest utilization metric for training clusters.
EVM · Earned Value Management
A method tracking project cost and schedule performance by comparing planned, earned and actual value.
EVPN · Ethernet VPN
The BGP-based control plane that tells a VXLAN fabric which endpoints live behind which switch; the standard overlay control plane in multi-tenant Ethernet designs.
Expert parallelism · EP
Distributing a Mixture-of-Experts model's experts across GPUs, routing tokens to whichever GPU holds the chosen expert.
Export controls
Government restrictions (administered by BIS) limiting which advanced AI chips can be sold to which countries.
Facility water · FWS
The building's water loop that ultimately rejects heat to the outdoors, kept separate from the clean chip-cooling loop.
FAT · Factory Acceptance Test
Testing equipment at the factory before shipment to confirm it meets specification, the first commissioning step.
Fat-tree
A multi-tier leaf/spine-derived network topology with increasing aggregate bandwidth toward the root. Oversubscription, path diversity, routing, link speeds, and failure state determine whether a particular deployment is nonblocking; the topology name alone does not.
Fault tolerance
The ability to withstand any single equipment failure without disrupting IT load, the defining trait of Tier IV.
FEC · Forward Error Correction
Encoding that lets the receiver fix transmission errors without retransmission, essential at high link speeds.
FedRAMP
The US government program standardizing security authorization for cloud services used by federal agencies.
FEOC · Foreign Entity of Concern
A designation restricting subsidies or participation for entities tied to certain adversary nations.
FERC · Federal Energy Regulatory Commission
The US agency regulating interstate electricity transmission, wholesale markets and grid interconnection rules.
FLAP-D
Europe's primary data-center markets: Frankfurt, London, Amsterdam, Paris and Dublin.
Floor loading
The structural weight a floor can bear; dense liquid-cooled racks can exceed limits and need reinforced slabs.
FLOPS · Floating-Point Operations Per Second
The standard measure of compute throughput; AI clusters are rated in petaFLOPS and exaFLOPS.
FM Global
A property insurer whose loss-prevention data sheets and equipment approvals act as a de facto design code, gating cooling-fluid, battery, and fire choices before coverage binds.
FMEA · Failure Mode and Effects Analysis
The discipline of enumerating how each component can fail, what the effect is, and what mitigates it — a commissioning and operations staple.
FOAK · First Of A Kind
The first deployment of a novel technology, carrying higher cost and risk than later, proven units.
Fork
A point where a design or program path splits into mutually exclusive options that must be chosen between.
FP16 · 16-bit floating point
A half-precision number format; it is the usual reference rung of the precision ladder, since each step down from it roughly doubles a chip's headline throughput.
FP32 · 32-bit floating point
Single-precision format, the top of the precision ladder; mixed-precision training keeps master weights and optimizer moments here even when the math runs lower.
FP4
A 4-bit floating-point format, native in Blackwell-class silicon, pushing inference throughput and density further.
FP6 · 6-bit floating point
An intermediate precision rung between FP8 and FP4, offered by some accelerators as a way to gain throughput without accepting the full quality risk of 4-bit.
FP8
An 8-bit floating-point format that roughly doubles throughput and halves memory versus 16-bit for AI workloads.
Free cooling · Economizer
Using cool outside air or water to reject heat without running mechanical chillers, saving energy in mild conditions.
FRU · Field-Replaceable Unit
A component designed to be swapped on-site, such as a power supply, fan or drive, without sending the whole system back.
FSDP · Fully Sharded Data Parallel
A training method that shards model parameters, gradients and optimizer state across GPUs to fit larger models.
FTR · Financial Transmission Right
The congestion hedge used in eastern US markets, functionally the counterpart to a congestion revenue right for managing basis between node and hub.
Gang scheduling
Scheduling all the GPUs a distributed job needs at once, so it either starts fully or waits, avoiding partial deadlock.
GB200
An NVIDIA Blackwell-generation system pairing two GPUs with a Grace CPU, deployed in the NVL72 rack-scale design.
GB300
NVIDIA's Blackwell Ultra rack generation (Grace CPU plus Blackwell Ultra GPUs) in NVL72 form; the density step between GB200 and the Rubin era.
GDPR · General Data Protection Regulation
The EU's comprehensive data-protection law governing how personal data is collected, processed and transferred.
GEMM · General Matrix Multiply
The dense matrix-multiply operation at the heart of neural-network compute and GPU benchmarking.
Genset
An engine-driven generator set providing backup or primary on-site power, typically diesel or natural gas.
Geo-redundancy
Replicating systems or data across geographically separate sites so one site's loss does not cause an outage.
GHG Protocol · Greenhouse Gas Protocol
The accounting standard defining emissions scopes and the location- versus market-based methods; its Scope 2 revision is moving claims toward hourly matching and deliverability.
GIS · Gas-insulated switchgear
Insulated-enclosure switchgear that packs high-voltage switching into a fraction of the air-insulated footprint, at high cost and some of the longest lead times on a project.
Glycol
Glycol is a family of heat-transfer-fluid base chemicals, not a complete coolant specification. A label such as PG25 states concentration only; base fluid, percentage basis, inhibitor package, water quality, materials compatibility, operating envelope, and equipment warranty must be named and approved.
GMP · Guaranteed Maximum Price
A contract capping the owner's cost; the contractor absorbs overruns above the agreed maximum.
GNSS · Global Navigation Satellite System
The satellite time reference a grandmaster clock disciplines to; because it can be jammed, spoofed, or lost, holdover oscillator quality is the real timing-plane design decision.
Golden image
A standardized, validated base system image cloned to every node for consistent, drift-free deployment.
Goodput
Useful work delivered per unit time after subtracting failed, restarted or stale work; the metric that actually matters.
GPAI · General-Purpose AI
The EU AI Act category covering foundation-model providers, carrying copyright-policy and training-data-summary obligations that reach anyone placing such a model on the EU market.
GPFS · General Parallel File System
IBM's mature parallel file system, shipped as Storage Scale, notable for distributing metadata across nodes rather than through a single metadata server and for broad multiprotocol support.
GPU · Graphics Processing Unit
A massively parallel processor that became the workhorse of AI training and inference.
GPU:CPU ratio
The number of GPUs per CPU in a node; AI servers skew heavily toward GPUs, reshaping system balance.
GPUDirect Storage · GDS
NVIDIA technology moving data directly from storage into GPU memory, bypassing the CPU bounce buffer.
Grace
NVIDIA's Arm-based server CPU, paired tightly with its GPUs over a coherent link in superchip designs.
Gray space
The back-of-house area housing power, cooling and infrastructure equipment that supports the white space.
Greenfield
A project built from scratch on undeveloped land, contrasted with brownfield reuse of an existing site.
Grid-forming inverter
An inverter that actively sets grid voltage and frequency, providing stability that conventional follow-the-grid inverters cannot.
GSU · Generator Step-Up transformer
Transformer that raises on-site generator output to grid or distribution voltage for behind-the-meter power.
H100
NVIDIA's Hopper-generation data-center GPU; the chip behind most of the 2023–2024 buildout wave and the baseline many density and power comparisons start from.
H200
NVIDIA's Hopper-generation GPU, a memory-upgraded H100 that serves as the common baseline generation in fleet, confidential-computing, and cross-vendor performance comparisons.
HBM · High-Bandwidth Memory
Stacked DRAM mounted on-package with the accelerator; the bandwidth and capacity ceiling and the key supply bottleneck.
HBM3E
An enhanced generation of high-bandwidth memory shipping in 2024-2025 accelerators, faster and denser than HBM3.
HBM4
The next high-bandwidth-memory generation with wider interfaces and a logic base die, targeting late-2020s accelerators.
HBOM · Hardware Bill of Materials
An inventory of the components in a hardware product, the hardware analog of an SBOM for supply-chain assurance.
HDD · hard disk drive
Rotating magnetic disk storage, now confined to the cold capacity tier and notable operationally because it cannot be immersion-cooled and is sanitized differently than flash.
Heat reuse
Capturing data-center waste heat to warm buildings or feed district heating, improving total energy use and ERF.
HGX
NVIDIA's baseboard reference platform integrating 8 GPUs with NVLink, the building block for many AI servers.
Hot spare
A spare node kept ready to swap in instantly when a failure occurs, minimizing interruption to a running job.
HPC · High-Performance Computing
The scientific supercomputing tradition AI clusters descend from — gang scheduling, low-latency fabrics, parallel filesystems — now repurposed for training at scale.
HSM · Hardware Security Module
A tamper-resistant device that generates, stores and uses cryptographic keys, the anchor of key custody.
HV · High voltage
The transmission-class voltage at the point of interconnection and the GSU transformer; HV breakers and switchgear are among the longest-lead items on a campus schedule.
HV transformer · High-Voltage transformer
Large transformer stepping transmission-level voltage down to site distribution; a long-lead item that often gates schedules.
Hybrid bonding
A copper-to-copper die-stacking technique with far finer, denser connections than solder microbumps.
Hyperscaler
A giant cloud and platform operator (AWS, Microsoft, Google, Meta) building data centers at global scale.
IaC · Infrastructure as Code
Managing infrastructure through version-controlled configuration files rather than manual setup, for repeatability.
IBC · International Building Code
The model building code adopted by most US jurisdictions; with the ASCE loading standard it sets the structural, seismic, and life-safety baseline the AHJ enforces.
ICI · Inter-Chip Interconnect
Google's TPU scale-up link, wiring chips into a torus instead of an all-to-all crossbar, so coherent domain size comes from topology and optical switching rather than a central switch tier.
IEC · International Electrotechnical Commission
The international electrotechnical standards body; its installation and equipment series govern designs outside North America, where the NEC and ANSI stack does not apply.
IEC 62443
The international standard series for cybersecurity of industrial automation and control systems.
IEEE · Institute of Electrical and Electronics Engineers
Standards organization whose electrical, switchgear, and media-sanitization specifications are written into data-center design bases and contracts alongside ANSI and NFPA.
IMEX · Internode Memory Exchange
The NVIDIA service and driver-level access control letting GPUs across a multi-node NVLink domain address each other's memory; the scheduler must provision and tear down its channels per job.
Immersion cooling
Submerging servers in a non-conductive dielectric fluid that carries heat away, in single-phase or boiling two-phase form.
IMS · Integrated Master Schedule
The master project schedule linking all tasks and dependencies, from which the critical path is derived.
Inferentia
Amazon's custom AI inference accelerator, optimized for cost-efficient serving of models.
InfiniBand · IB
A low-latency lossless fabric with native RDMA, the historical default for non-blocking AI training back-ends.
Interconnection queue
The utility waitlist for connecting new large loads or generation to the grid; multi-year waits dominate AI siting.
Interposer
The silicon or organic layer carrying dense wiring between logic and memory in a 2.5D package.
IOPS · Input/Output Operations Per Second
The rate of read/write operations a storage system handles; metadata-heavy AI loads are often IOPS-bound, not bandwidth-bound.
IPMI · Intelligent Platform Management Interface
The legacy out-of-band server-management protocol with weak authentication and a long vulnerability history; fleets still leaning on it carry a management-plane security liability.
IRR · Internal Rate of Return
The discount rate at which an investment's net present value is zero, a headline measure of project return.
IRU · Indefeasible Right of Use
A long-term lease of unlit fiber strands you light yourself, giving control of bandwidth, routing, and latency path at the cost of a far longer delivery timeline than buying lit capacity.
ISMS · Information Security Management System
The documented, audited security-management program that ISO 27001 certifies; maintaining it is a continuing obligation with recurring surveillance audits, not a one-time exercise.
ISO · Independent System Operator
An operator managing grid reliability and power markets in a region, often used interchangeably with RTO.
ISO 22237
The international data-center facility standard series, the global counterpart to Europe's EN 50600.
ISO 27001
The international standard for an information security management system (ISMS), a baseline enterprise security certification.
ISO 42001
The international standard for an AI management system (AIMS), governing responsible development and operation of AI.
ISSB · International Sustainability Standards Board
The body issuing the global baseline for sustainability disclosure, adopted into many jurisdictions' reporting rules alongside the EU regime.
IST · Integrated Systems Test
The final commissioning level (L5) that proves all facility systems work together under simulated full load.
ITAD · IT asset disposition
The certified chain-of-custody process for sanitizing, reselling, or destroying retired equipment; it is a supply-chain link with its own insider risk, not a disposal afterthought.
ITUE · IT-power Usage Effectiveness
An efficiency metric pushing the boundary inside the server to capture fan, VRM and PSU losses.
JAX
A machine-learning framework compiled through XLA, dominant on TPUs and supported elsewhere; picking it ties your model code to that compiler's toolchain.
JCT · job completion time
Wall-clock time from training job start to a finished model, the north-star fabric metric because it folds in tail latency, congestion, goodput, failures, and restarts.
Jevons paradox
The principle that efficiency gains can raise total consumption; cheaper AI inference spurs far more demand, not less.
Junction temperature · Tj
The temperature of the actual transistors inside a chip; exceeding its max forces throttling or damage.
KMS · Key Management Service
A KMS manages cryptographic keys. Destroying a key can sanitize data only when the applicable NIST method permits crypto erase and encryption/key generation, key custody and destruction, media coverage, verification, records, and final disposition satisfy the system's policy; encryption alone is not automatic compliance.
KPI · Key Performance Indicator
A tracked operating metric; data-center KPI families cover efficiency (PUE, WUE), reliability, and utilization.
Kubernetes · K8s
The dominant open-source container orchestration platform, increasingly used to schedule AI inference and training.
KV cache · Key-Value cache
Stored attention tensors reused during decode; its size grows with context and concurrency, dominating inference memory.
Kyber
NVIDIA's rack-scale platform for the Rubin Ultra generation, scaling NVLink domains to hundreds of GPUs.
L2A · Liquid-to-air
Trade shorthand for a closed-loop cold-plate system rejecting heat into the existing air plant: the drop-in retrofit path, bought at the price of a hard density ceiling.
L2L · Liquid-to-liquid
Trade shorthand for full direct-to-chip cooling that rejects into a facility water loop through CDUs: the highest-density retrofit path and the most disruptive to a live hall.
Lakehouse
A data architecture (Iceberg, Delta, Hudi) adding database-like transactions and schema to cheap object storage.
LBNL · Lawrence Berkeley National Laboratory
The US national lab whose interconnection-queue and data-center energy studies are the standard public reference for grid wait times and load growth.
LCOE · Levelized Cost of Energy
The per-MWh lifetime cost used to rank generation options; on a build gated by time-to-megawatt it is the wrong ranking variable used alone.
Lemon node
A subtly defective node that passes basic checks but repeatedly degrades jobs, found by lemon-node detection.
LFP · Lithium iron phosphate
The lithium chemistry favored for facility storage and rack backup because it runs cooler and tolerates abuse better than nickel-based cells, easing the thermal-runaway and fire case.
LGIA · Large Generator Interconnection Agreement
The contract governing how a large facility or generator connects to the transmission grid.
Lights-out operations
Running a facility with minimal on-site staff, relying on remote management and automation.
LLMflation
The rapid collapse in the cost to serve a given level of AI capability as models and hardware improve.
LMP · Locational Marginal Price
The price of electricity at a specific grid node, reflecting local supply, demand and congestion.
LNG · Liquefied natural gas
Cryogenically liquefied gas trucked and stored on site; it backstops pipeline curtailment and can bridge on-site generation before a firm gas lateral exists.
Load step · Power transient
A sudden synchronized swing in GPU draw across thousands of chips that stresses the power chain in milliseconds.
Long-lead equipment
Items like transformers, switchgear and chillers whose long procurement times often gate the schedule.
LoRA · Low-Rank Adaptation
A fine-tuning method that trains small low-rank adapter matrices instead of all weights, slashing cost and memory.
Lose a nine
Reliability shorthand for multiplying the unavailability or downtime fraction by 10 over a stated observation period and service boundary; it is not a context-free availability-percentage subtraction.
LOSF · Lots Of Small Files
A workload pattern of huge numbers of tiny files that stresses storage metadata far more than raw bandwidth.
LOTO · Lockout/tagout
The disciplined de-energization and isolation regime that makes work on energized-capable plant survivable; supervising it is a verified competency, not a paperwork step.
LPDDR5X · low-power double data rate 5X (SDRAM)
Low-power DRAM soldered into coherent CPU-GPU superchip packages; it buys bandwidth and efficiency over socketed DDR5 at the cost of field-serviceability and capacity choice.
LPO · Linear Pluggable Optics
Pluggable optics with the DSP removed so the host ASIC's SerDes drives the optic directly; it cuts module power sharply but makes you the owner of the entire electrical path and its interop risk.
LPS · Lightning protection system
The air terminals, down conductors, and low-impedance earth connection on the building envelope that give a strike a path around the electronics rather than through them.
LV · Low voltage
The sub-kilovolt class of final distribution boards, PDUs, and server inputs, where conductor sizing, fault current, and copper cost dominate the design.
Maia
Microsoft's custom AI accelerator chip, part of its in-house silicon program for Azure AI.
Makeup water
Fresh water added to a cooling system to replace what evaporation, drift and blowdown remove.
MBU · Model Bandwidth Utilization
Achieved memory bandwidth over peak for memory-bound decode inference; the MBU is the MFU analog when HBM-bound.
MCTP · Management Component Transport Protocol
The in-platform transport carrying management traffic between the BMC and components; it is how out-of-band firmware updates reach GPUs and NICs.
Measured boot · Attestation
Measured boot records cryptographic measurements of boot components, typically into TPM PCRs, for later attestation. A measurement is not a trust decision: verification also needs an authenticated quote, approved reference values/policy, freshness, and a defined TCB and remediation path.
MEC · Multi-access Edge Computing
Placing compute near users at the network edge (e.g. telco sites) to cut latency for real-time AI services.
Memory-bandwidth-bound
A workload limited by how fast data moves from memory rather than by compute; typical of inference decode.
MEP · Mechanical, Electrical and Plumbing
The building-services scope of a construction project; in data centers the MEP package dominates cost, complexity, and schedule risk.
Merchant
Revenue or power sold on the open market without a long-term contract, carrying price risk versus contracted supply.
MFA · Multi-Factor Authentication
Requiring more than one credential factor; escalating it at each physical and logical zone boundary is the baseline control every compliance regime credits at once.
MFU · Model FLOPs Utilization
Achieved FLOPs divided by peak FLOPs in a training run; 35-55% is good at scale, eroded by collectives and stragglers.
MGX
NVIDIA's modular server reference architecture letting partners build varied GPU systems from common building blocks.
MI300X
AMD's first broadly deployed AI GPU and the canonical case study in the gap between paper FLOPS and realized throughput once the software stack is accounted for.
MI355X
AMD's MI350-series AI GPU, the generation whose ROCm support first made a credible multi-thousand-GPU second source, and a clean illustration of precision-driven FLOPS inflation.
MI455X
AMD's MI450-series AI GPU, the accelerator inside the Helios rack-scale system and the basis of the open-standards alternative built on OCP, UALink, and Ultra Ethernet.
MIG · Multi-Instance GPU
An NVIDIA feature partitioning one GPU into isolated instances so multiple workloads share it securely.
MISO · Midcontinent Independent System Operator
The regional grid operator covering the US midcontinent, one of the balancing authorities where large-load headroom and interconnection rules differ materially from neighbors.
MLPerf
The industry-standard benchmark suite for AI training and inference performance; the closest thing to apples-to-apples numbers across accelerator vendors.
MMLU · Massive Multitask Language Understanding
A broad knowledge benchmark used as the quality yardstick when judging whether a lower-precision training or quantization recipe has degraded a model.
MOC · Management of Change
The facility-side change-control process that sorts every proposed change into a risk tier and gates the tiers that can disturb live load.
Modbus
The lingua franca of industrial controllers; unauthenticated by design, so anything that can reach the network can command a chiller or CDU without exploiting a vulnerability.
MoE · Mixture of Experts
A sparse model that activates only a subset of expert sub-networks per token, cutting compute per token at scale.
MOP · Method of Procedure
The step-by-step script for a specific, planned, often one-time intervention on live infrastructure; it is what makes non-routine work auditable rather than improvised.
MPO · Multi-fiber Push-On
The multi-fiber connector terminating parallel-lane optics; its base-N grouping and polarity scheme decide whether the fiber plant survives the next lane-rate generation or has to be re-pulled.
MPS · Multi-Process Service
NVIDIA's mechanism for letting multiple processes share one GPU context concurrently; it improves packing for same-trust workloads but is a soft boundary, not a security wall.
MRC · Multipath Reliable Connection
An open RoCEv2 extension adding packet spray, trimming, and selective acknowledgment; already proven at frontier training scale, but training-first in scope rather than a general-purpose transport.
Mt Diablo
OCP/Meta's Mt Diablo rack-power architecture uses a bipolar ±400 VDC bus about a midpoint (800 V rail-to-rail). It is distinct from NVIDIA's announced unipolar 800 VDC rack architecture; do not merge the two profiles.
MTBF · Mean Time Between Failures
The average operating time between failures of a component or system, a core reliability input.
MTIA · Meta Training and Inference Accelerator
Meta's custom AI accelerator family, built to run its recommendation and language workloads more cheaply than GPUs.
MTTI · Mean Time To Interruption
The average time a large training job runs before something interrupts it, a key scaling-reliability metric.
MTTR · Mean Time To Repair
The average time to restore a failed component to service, a core driver of overall availability.
MV · Medium voltage
The utility-to-facility distribution voltage class feeding site transformers and switchgear; the level where lead times, arc-flash energy, and gear cost step up sharply.
MVA · Megavolt-Ampere
Unit of apparent power used for transformer nameplate rating. Switchgear must instead be specified by voltage, continuous current, short-circuit interrupting/withstand duty, insulation/BIL, topology, and applicable standard; do not collapse switchgear duty to MVA.
MW · Megawatt
Unit of real power; in AI data centers it has become the de facto unit of compute capacity and grid demand.
N+1
A component group or capacity arrangement with N units required for the declared design load plus one spare unit. N+1 describes provisioned capacity, not guaranteed tolerance of every failure or maintenance event or an availability outcome.
NCCL · NVIDIA Collective Communications Library
NVIDIA's library implementing optimized multi-GPU collective operations like all-reduce over NVLink and the fabric.
NDE · Non-destructive examination
Radiographic, ultrasonic, or penetrant inspection of welds; the piping code you adopt fixes how much NDE and documentation you owe for the life of the facility.
Neocloud
A new breed of GPU-focused cloud provider (CoreWeave, Lambda and peers) renting AI compute outside the big hyperscalers.
NEPA · National Environmental Policy Act
The US law requiring federal projects to assess environmental impacts, a potential permitting gate for some sites.
NERC
The North American body setting and enforcing mandatory grid reliability standards, including critical-infrastructure rules.
NERC CIP · Critical Infrastructure Protection
The mandatory, audited cyber and physical security standards for bulk-electric-system assets, which reach a campus through its interconnection and any grid-facing equipment it owns.
NeuronLink
AWS's scale-up interconnect for Trainium, paired with a switched all-to-all topology inside an UltraServer to give uniform any-to-any bandwidth across the coherent domain.
NFPA · National Fire Protection Association
The body publishing the electrical and fire codes that gate data-hall electrical design, lithium-battery placement, and suppression schemes.
NIC · Network Interface Card
The server's network adapter. AI nodes carry multiple RDMA-capable high-speed NICs — often one per GPU — as the on-ramp to the scale-out fabric.
NIST · National Institute of Standards and Technology
The US institute whose control catalogs and media-sanitization guidance underpin federal compliance regimes and end-of-life handling of fleet storage.
NIST SP 800-53
The US catalog of security and privacy controls for federal information systems, a foundation for many compliance regimes.
NIXL · NVIDIA Inference Xfer Library
The de facto library for moving KV cache across memory, network, and storage tiers, which is what makes prefill/decode disaggregation practical by turning KV transfer into a scheduled object.
Non-blocking
A network that can carry full bandwidth between all node pairs simultaneously with no internal contention (1:1).
NOx · Nitrogen oxides
The combustion pollutant that sets the air-permitting envelope for on-site generation; the cheapest NOx is the NOx never combusted, the next cheapest is abated by design.
NPDES
The US permit program regulating pollutant discharges to surface waters, governing cooling-water blowdown.
NPSH · Net positive suction head
The suction-side pressure margin a pump needs to avoid cavitation; available must beat required at the worst-case low-pressure node in every redundancy state.
NPV · Net Present Value
The present value of future cash flows minus the investment, the core go/no-go metric for capital projects.
NRE · Non-Recurring Engineering
The one-time design cost of a custom chip or system, absorbed up front for a lower cost per unit — and unrecoverable if the volume never arrives.
NREL · National Renewable Energy Laboratory
The US national lab for renewable-energy and efficiency research, a common source for water- and energy-intensity reference figures.
NSPS · New Source Performance Standards
Federal emission standards for new combustion sources; their tightening is closing the 'temporary turbine' classification that made bridge generation quick to permit.
NTU · Number of transfer units
The heat-exchanger sizing measure tying effectiveness to surface area and conductance; more NTU buys a tighter approach temperature, which buys warmer usable facility water.
NUMA · non-uniform memory access
The property that a CPU socket reaches some memory and devices faster than others; misaligned GPU-NIC-socket affinity silently drags down collective performance across an entire job.
NVFP4
NVIDIA's 4-bit microscaling format, using small value blocks with a floating-point scale; the choice between it and MXFP4 is baked into silicon and into checkpoint portability.
NVL72
An NVIDIA rack connecting 72 Blackwell GPUs into one NVLink domain that behaves as a single huge accelerator.
NVLink
NVIDIA's high-bandwidth GPU-to-GPU interconnect for tying many GPUs into one memory-coherent scale-up domain.
NVMe · Non-Volatile Memory Express
The high-speed protocol for SSDs over PCIe, the storage interface standard in modern AI servers.
NVML · NVIDIA Management Library
The in-band API exposing per-GPU temperature, power, clocks, and error state; it is the bottom layer that DCGM and fleet observability stacks are built on.
NVSwitch
NVIDIA's switch chip that fully connects all GPUs in a scale-up domain and can do in-network reductions.
Object storage
Storage that keeps data as objects in a flat namespace accessed by API (e.g. S3), the backbone for AI datasets.
OCP · Open Compute Project
An industry community that open-sources data-center hardware designs for racks, power, cooling and security.
OEM · Original Equipment Manufacturer
A systems vendor that turns chipmakers' reference designs into deployable servers and racks; the integration, support, and warranty point between silicon and operators.
OFE · Owner-Furnished Equipment
Gear the owner buys directly and free-issues to the contractor; it strips contractor margin and controls the order date but transfers schedule and interface risk back.
Off-gas detection
Sensing the gases a failing battery cell vents, an early-warning trigger before thermal runaway and fire.
OOB · Out-of-Band
A management path separate from the production network (BMC, serial console, management fabric) that keeps control reachable when the data plane is down.
Opex · Operating Expenditure
Ongoing running costs such as power, water, staff and maintenance, expensed as incurred.
OPR · Owner's Project Requirements
The owner's intent expressed in measurable terms; every commissioning script must trace back to it, so a vague OPR guarantees an unfalsifiable acceptance test.
Optimizer state
The extra per-parameter data an optimizer like Adam keeps (momentum, variance), often doubling or tripling memory needs.
ORR · Operational Readiness Review
A formal gate confirming a facility or cluster is ready to safely take on production load.
ORV3 · Open Rack V3
The Open Compute Project's third-generation rack standard, defining power, busbar and form-factor for open hardware.
OS2
A low-attenuation single-mode optical-cabling category. Achievable reach and upgrade reuse still depend on link budget, wavelength, optic class, connectors/splices, dispersion, fiber condition, and qualification—not a generation-agnostic or transceiver-only guarantee.
OSAT · Outsourced Semiconductor Assembly and Test
Companies that package and test chips after fabrication, a key and capacity-constrained step for advanced AI silicon.
OSFP · Octal Small Form-factor Pluggable
Pluggable transceiver form factor with a larger thermal envelope than QSFP-DD, favored for the hottest high-rate AI links; the cage is soldered to the switch and fixed for its life.
OT · Operational Technology
The control systems running physical infrastructure (power, cooling, building management), a growing cyber-attack surface.
Oversubscription
Provisioning less network bandwidth than full non-blocking would require; 1:1 is non-blocking, 3:1 is oversubscribed.
P50 / P90
Confidence levels for a schedule or estimate: P50 is the median outcome, P90 the value met 90% of the time.
PagedAttention
A vLLM technique that manages the KV cache in fixed pages like virtual memory, cutting waste and fragmentation.
PAM · Privileged Access Management
Tooling that brokers administrative access through vaulted, session-recorded, just-in-time grants — what makes no-standing-access real for management-plane and OT jump hosts.
PAM4 · Pulse Amplitude Modulation, 4-level
Four-level signaling that carries two bits per symbol to double lane rate, at the cost of a much smaller eye that makes forward error correction and transmitter margin mandatory.
Parallel file system
A storage system (Lustre, GPFS, WEKA, VAST) serving many clients at once with high aggregate bandwidth for AI clusters.
PAUSE
The Ethernet back-pressure frame telling an upstream link to stop sending; because it propagates hop by hop, it is the mechanism behind head-of-line blocking, pause storms, and fabric deadlock.
PCB · Printed Circuit Board · Parts 4, 7, 8
The laminated board that carries and connects a system's components; at AI power and signal densities its trace losses and midplane design become first-order engineering constraints.
PCB · Polychlorinated Biphenyls · Parts 3, 14
Toxic compounds found in legacy transformer dielectric oils and older sites; their testing and disposal obligations surface in site due diligence and decommissioning.
PCIe · Peripheral Component Interconnect Express
The standard high-speed bus connecting CPUs, GPUs, NICs and storage inside a server.
PDU · Power Distribution Unit
Equipment that distributes electrical power to racks; rack PDUs are the metered strips feeding individual servers.
PE · Protective earth
The earthing conductor carried separately from neutral to every load, giving fault current a defined low-impedance return so protective devices clear and enclosures stay safe to touch.
PED · Pressure Equipment Directive
EU law covering pressure equipment above a threshold; on European sites it, not a voluntary code, fixes weld QA, inspection extent, and pressure testing for cooling loops.
PEFT · Parameter-Efficient Fine-Tuning
A family of methods (like LoRA) that adapt large models by training only a tiny fraction of parameters.
PFAS
Persistent fluorinated 'forever chemicals' found in some dielectric coolants, raising environmental and regulatory concern.
PFC · Priority Flow Control
An Ethernet mechanism that pauses traffic to prevent packet loss, making RoCE lossless but risking head-of-line blocking.
PG25 · 25% propylene glycol
PG25 denotes a nominal 25% propylene-glycol concentration; it is not a universal direct-to-chip default or a complete fluid specification. Name the product/base fluid, percentage basis, inhibitor package, water quality, materials compatibility, operating envelope, maintenance controls, and equipment warranty.
Phase gate · Stage gate
A go/no-go decision point between project phases where deliverables are reviewed and capital is released.
Phase I ESA · Environmental Site Assessment
A desktop and walkover review of a site's environmental history to flag contamination risk before purchase.
PHC · PTP Hardware Clock
The clock inside a NIC that a PTP daemon disciplines to the grandmaster and that the host system clock then follows; fleet-wide offset from master is what commissioning must prove.
PHY · Physical Layer
The physical-layer interface of a link, whether an Ethernet port or a die-to-die crossing; its ecosystem maturity, power, and silicon area often decide build-now versus wait-for-new-silicon.
PID · Proportional-integral-derivative
The classical feedback loop running pumps, valves, and CDUs; its tuning sets overshoot on a synchronized load step, and it stays as the fallback beneath any learned controller.
PII · Personally Identifiable Information
Data identifying a person; cheapest handled by detection and redaction before it enters a training corpus, because removing it after training is far harder.
PILOT · Payment In Lieu Of Taxes
A negotiated payment a data center makes instead of standard property taxes, often part of siting incentives.
Pipeline bubble
Idle GPU time at the start and end of pipeline-parallel execution while the pipeline fills and drains.
Pipeline parallelism · PP
Splitting a model's layers across GPU groups that process different micro-batches in an assembly line.
PJM
The largest US regional grid operator, covering the mid-Atlantic and a major data-center hub with long queues.
PKey · Partition Key
InfiniBand partition key used to segment a back-end fabric; it is a configuration boundary enforced by the subnet manager and adapters, not cryptographic isolation.
PLC · Programmable logic controller
The industrial controller executing chiller, CDU, and switchgear I/O beneath the BMS and EPMS; it usually speaks unauthenticated protocols, which makes it prime OT attack surface.
PLDM · Platform Level Data Model
The standardized component firmware-update and monitoring model carried over MCTP, making signed, host-independent firmware updates tractable across a mixed fleet.
Point of interconnection · POI
The physical point where a facility's electrical system connects to the utility grid.
PoP · Point of Presence
A network access location where a provider's infrastructure meets users or other networks.
POSIX · Portable Operating System Interface
The file-system semantics standard that gives a byte-addressable, strictly consistent namespace; choosing a POSIX parallel file system over an object-native store is a primary storage fork.
Post-training
The fine-tuning and alignment stages (SFT, RLHF) after pretraining that shape a model's behavior and usefulness.
Power capping
Limiting how much power chips or racks can draw to stay within facility limits, at the cost of some performance.
Power factor
Ratio of real power (kW) to apparent power (kVA); a low power factor wastes capacity through reactive current.
Power oversubscription
Provisioning more IT than the power can sustain at full draw, relying on workloads rarely peaking together.
Powered shell
A building delivered with power and core infrastructure in place but without IT fit-out, ready for a tenant to finish.
PPA · Power Purchase Agreement
Long-term contract to buy electricity (often renewable) at a set price, used to secure and green a site's power.
PPE · Personal protective equipment
Arc-rated clothing and gear selected from the incident-energy study, which means a relay setting that has drifted from the model can invalidate the label a technician is trusting.
PQ · Power quality
Harmonics, sags, swells, and transients on the bus; specifying PQ-capable metering at the right tiers is what turns a metering hierarchy into usable diagnostic evidence.
PQC · Post-Quantum Cryptography
Cryptography designed to resist quantum attack; assets whose attestation and signing keys must outlive the transition should root trust in quantum-resilient algorithms from the start.
Preemption
Pausing or evicting a lower-priority job to free resources for a higher-priority one in a shared cluster.
Prefill
The compute-heavy phase that processes an inference prompt and builds its KV cache before generation begins.
Pretraining
The initial, compute-heavy phase that trains a model on vast unlabeled data to learn general capabilities.
Provenance register
A record linking assertions about a component or artifact to sources, origin, transformations, ownership, and custody. Provenance supports auditability but does not alone prove authenticity or integrity; signatures, verification, controls, and evidence quality remain necessary.
PSD · Prevention of Significant Deterioration
The permitting path a major new emitting source triggers in an attainment area, requiring best-available control technology and dispersion modeling — a year-plus schedule item.
PSM · Process Safety Management
The US occupational-safety regime for covered flammable-inventory processes; on-site gas or fuel above threshold pulls its full element set into scope before first fuel.
PSU · Power Supply Unit
The component that converts incoming AC or DC into the regulated voltages a server or GPU node consumes.
PTP · Precision Time Protocol
IEEE-1588 time distribution using hardware timestamping in NIC and switch silicon; treat its offset bound under saturated load as a measured acceptance gate, not a configuration step.
PUE · Power Usage Effectiveness
Total facility power divided by IT power; the headline efficiency ratio where 1.0 is perfect and AI halls target ~1.1-1.2.
Purdue model
A reference architecture that layers and segments industrial control networks to contain cyber threats.
PXE boot · Preboot Execution Environment
Booting a server over the network to load its OS image, the basis of automated bare-metal provisioning.
PyTorch
The dominant open-source machine-learning framework and the one every accelerator vendor's software stack is tuned against first, which is why second-source silicon is judged by its PyTorch support.
QLC · Quad-Level Cell
Flash storing four bits per cell, offering high density and low cost at the expense of write endurance and speed.
QoS · Quality of Service
Mechanisms that give some traffic or jobs priority over others — network traffic classes, scheduler priority tiers, preemption rules.
Quantization
Using lower numerical precision (FP8, FP4, INT8) to cut memory and boost throughput, trading some accuracy for speed.
Quick-disconnect · QD
A self-sealing coupling that connects or separates a liquid line without leaks, used to service cooled racks.
RAG · retrieval-augmented generation
Serving pattern that retrieves external context and prepends it to the prompt, which makes traffic prefix-heavy and therefore highly sensitive to KV-cache reuse and routing.
Rail-optimized
A topology pinning each GPU's NIC to a dedicated switch 'rail' for collision-free, non-blocking training collectives.
RAM · random-access memory
Volatile working memory; in AI nodes the term usually means host RAM, which stages data-loader batches and holds the fastest checkpoint tier alongside the GPU's own HBM.
RBAC · Role-Based Access Control
Granting permissions by role rather than by individual; the substrate of tenant isolation and quota enforcement in cluster control planes.
RBD · Reliability Block Diagram
A modeling technique that maps components as a network of blocks to compute overall system availability.
RCCL · ROCm Communication Collectives Library
AMD's collective-communication library and NCCL counterpart; its maturity, more than raw silicon, is what gates AMD's viability for tightly-coupled frontier training.
RDHx · Rear-Door Heat Exchanger
A liquid-cooled rack door removing ~50-100 kW without piping liquid to the chips; a brownfield step toward full DLC.
RDMA · Remote Direct Memory Access
Network transfers that move data directly between machines' memory, bypassing the CPU for low latency.
Reactive power
Power that oscillates between source and load without doing work, measured in VAR, that must be managed and corrected.
REC · Renewable Energy Certificate
A tradable certificate representing one MWh of renewable generation, used to claim clean-energy use.
Redfish
A modern REST API standard for remotely managing server and infrastructure hardware, succeeding legacy IPMI.
REF · Renewable Energy Factor
The share of a facility's energy supplied from renewable sources, a sustainability companion to PUE.
Reticle
The maximum area a lithography tool can pattern in one exposure (~858 mm2), capping how large a single die can be.
Reversibility
Designing a decision so it can be undone cheaply; reversible choices warrant less analysis than one-way doors.
Reward model
A model trained to score outputs, providing the reward signal that guides reinforcement-learning fine-tuning.
RFP · Request for Proposal
The competitive solicitation document, and the place where resilience and density claims must be pinned to testable criteria rather than data-sheet labels.
RFS · Ready For Service
The milestone at which a facility or capacity block is fully commissioned and available to take load.
RICE · Reciprocating internal combustion engine
Gas or diesel reciprocating engines used for backup and bridge generation: modular, fast to deploy, and flexible at part load, at the cost of emissions and unit count.
Ride-through
A system's ability to stay running through a brief power or cooling disturbance instead of tripping offline.
RIM · Reference Integrity Manifest
A signed reference of expected firmware measurements used to verify a device booted untampered code.
RL · Reinforcement Learning
Training that optimizes a model against a reward signal rather than labeled data; post-training RL mixes training and inference and loads clusters in spiky patterns.
RLHF · Reinforcement Learning from Human Feedback
Tuning a model using human preference rankings to make its outputs more helpful and aligned.
RMA · Return Merchandise Authorization
The process of returning failed hardware to a vendor for repair or replacement under warranty.
RoCE · RDMA over Converged Ethernet
RDMA over Ethernet. ECN signals congestion; it does not make the fabric lossless. The design must state queueing and congestion control plus any PFC, retransmission/recovery, routing, buffering, and failure-handling choices.
ROCm
AMD's open GPU-computing software stack, its counterpart to NVIDIA's CUDA ecosystem.
Rollout
In reinforcement learning, generating a trajectory of model actions and outcomes used to compute training rewards.
Root of trust · RoT
A hardware-anchored trusted base that verifies firmware and boot integrity before a system is trusted to run.
RTO · Regional Transmission Organization
An entity operating the grid and wholesale power market across multiple utilities in a region (e.g. PJM, ERCOT).
RTT · Round-Trip Time
The there-and-back network delay; it sets the interactive-inference latency budget and, between sites, decides whether a training run can remain synchronous at all.
Rubin
NVIDIA's GPU architecture generation following Blackwell, with Rubin Ultra pushing rack-scale density further.
S3 · Simple Storage Service
Amazon's object-storage service whose API has become the de facto standard for cloud object storage.
SBOM · Software Bill of Materials
A formal inventory of all components in a piece of software, used to track and respond to supply-chain risk.
SCADA · Supervisory Control and Data Acquisition
Industrial control software that monitors and operates a facility's physical systems like power and cooling.
Scale-out
Connecting many nodes into a looser cluster fabric (InfiniBand or Ethernet) to scale beyond one coherent domain.
Scale-up
Tightly coupling GPUs into one coherent high-bandwidth domain (NVLink class) that acts like a single large accelerator.
Scaling laws
Empirical relationships predicting how model quality improves with more compute, data and parameters.
Scope 1
Direct greenhouse-gas emissions from sources a company owns or controls, such as on-site generators.
Scope 2
Indirect emissions from the purchased electricity, heat or cooling a facility consumes.
Scope 3
All other indirect emissions in the value chain, including the embodied carbon of equipment and construction.
SCR · Selective catalytic reduction
The exhaust aftertreatment that strips NOx from engines and turbines; sizing it in at design time is far cheaper than retrofitting it under enforcement.
SDC · Silent Data Corruption
Errors that corrupt computation without any alert; at fleet scale they silently spoil training and must be hunted.
Secure boot
A boot process that cryptographically verifies each firmware and software stage before allowing it to run.
SemiAnalysis
An independent semiconductor and data-center research firm whose cost models, teardowns, and GPU-cloud ratings are widely used as industry reference points.
SerDes · Serializer/Deserializer
The circuit converting parallel data to a high-speed serial stream and back; lane rate sets link bandwidth.
SEV-SNP · Secure Encrypted Virtualization - Secure Nested Paging
AMD's confidential-VM technology, encrypting each guest's memory under a per-VM key held by an on-die secure processor; the AMD-side counterpart to Intel TDX.
SFT · Supervised Fine-Tuning
Adapting a pretrained model by training it on labeled example responses for a target behavior or domain.
SGD · stochastic gradient descent
The iterative optimization loop underlying model training; whether it synchronizes every step or only intermittently determines how far apart training sites can physically sit.
SGLang
An open-source LLM serving framework with fast structured generation and aggressive KV-cache reuse.
SHARP · Scalable Hierarchical Aggregation and Reduction Protocol
In-network computing that lets switches perform parts of a collective (such as summing gradients) so data crosses the fabric fewer times.
SiC · Silicon carbide
The wide-bandgap power semiconductor behind active front ends, solid-state transformers, and dense rectifiers; it switches faster and hotter than silicon, buying efficiency and footprint.
SIEM · Security Information and Event Management
The platform aggregating security telemetry for detection; AI campuses defeat naive deployments because fabric, GPU, facility, and physical-access signal volumes overwhelm conventional ingest.
SIS · Safety-instrumented system
An instrumented protection system designed to drive a process to a safe state and meet a specified safety-integrity requirement. Architecture may be hardwired or programmable; independence, override/bypass controls, diagnostics, proof testing, and lifecycle governance are design-specific, not universal.
SK hynix
One of the three suppliers able to make high-bandwidth memory at scale, so its stacking capacity and qualification schedule help set the ceiling on accelerator supply.
SKU · Stock Keeping Unit
A specific vendor part configuration; firmware defects are often SKU-specific, so a canary set has to span every SKU and supplier in the fleet.
SLA · Service Level Agreement
A contractual promise of service performance (uptime, latency) with penalties or credits if it is missed.
SLO · Service Level Objective
A target for a service metric such as latency or availability that an inference fleet is sized to meet.
Slurm
A widely used open-source workload manager and job scheduler for HPC and AI clusters.
SM · streaming multiprocessor
The GPU's fundamental compute block; SM occupancy is the utilization signal autoscalers read, and it misleads badly on decode work that is memory-bandwidth-bound rather than compute-bound.
SmartNIC
A network card with onboard processing that offloads packet, storage and security work from the server CPU.
SMI · System Management Interface
NVIDIA's management control surface for reading GPU telemetry and applying power caps and clock limits; with Redfish it is how workload-side power smoothing is actually enacted.
SMR · Small Modular Reactor
Factory-built nuclear reactor under ~300 MW, proposed as clean firm power for large AI campuses.
SNMP · Simple Network Management Protocol
The legacy polling protocol still carrying much facility and network telemetry; reconciling it with Modbus, BACnet, and Redfish is the cost of one observability pane.
SoC · system on chip
A complete computer integrated onto one die; every server contains several beyond the host CPU, and each carries its own firmware and therefore its own attack surface.
SOC 2
An audit report (Type II covers a period) attesting that an organization's security and availability controls work as described.
SOFC · Solid-Oxide Fuel Cell
A high-temperature fuel cell generating clean on-site electricity from gas or hydrogen for data-center power.
SOO · Sequence of Operations
The control-logic specification defining how the BMS, EPMS, and DCIM must behave in every normal, maintenance, and failure mode; functional testing is a line-by-line proof against it.
SOP · Standard Operating Procedure
The routine, repeatable procedure for steady-state operation, and the baseline against which any deviation or non-routine intervention has to be justified.
Sovereign AI
A nation's drive to own and control AI compute, data and models within its borders for security and autonomy.
SPD · Surge protective device
A surge-diverting device coordinated in stages from service entrance to equipment, keeping lightning and switching transients off sensitive electronics.
Spectrum-X
NVIDIA's Ethernet platform tuning RoCE for AI collectives with adaptive routing and congestion control.
Speculative decoding
Using a small draft model to guess several tokens that a large model verifies in parallel, speeding generation.
SPOF · Single Point of Failure
A component whose failure alone takes down the whole system; eliminating SPOFs is the goal of redundant design.
SPP · Southwest Power Pool
The regional grid operator spanning the central US plains, a market regime with its own study sequence, cost allocation, and curtailment obligations.
SPV · Special-Purpose Vehicle
A standalone, bankruptcy-remote legal entity created to own and finance a single project and ring-fence its risk.
SRAM · static random-access memory
Fast on-die memory used for caches and scratchpads, and the basis of HBM-free inference chips; it is also a distinct GPU failure mode that first appears as correctable-error storms.
SRv6 · Segment Routing over IPv6
Source routing where hosts stamp the path onto each packet and switches forward from static tables, moving failure re-pathing out of the fabric control plane and into software you own forever.
SSD · solid-state drive
Flash-based storage with no moving parts, the default medium for the node-local scratch and checkpoint tiers where sustained per-drive bandwidth matters more than capacity cost.
SST · Solid-State Transformer
Solid-state transformer/power-electronics conversion from medium voltage toward a rack DC bus. About 98% is a named 400 kW prototype result; about 99% is a target, not a general shipping-product efficiency. Preserve topology, load point, voltage, and measurement boundary.
Straggler
A node running slower than its peers that holds up a synchronized collective and drags down whole-job throughput.
Stranded capacity
Provisioned power, cooling or space that cannot be used because a different resource is the binding constraint.
Stranded power
Generation or interconnection capacity that exists but cannot reach load due to transmission or siting limits.
STS · Static Transfer Switch
A solid-state switch that instantly transfers a load between two power sources without interruption.
SU · Scalable Unit
A repeatable build block (a defined MW + GPU + cooling + fabric increment) that capacity ramps are composed of.
Substation
The facility transforming and switching power between transmission and the site, a major long-lead build item.
Super-load · Heavy-haul
An oversized, very heavy shipment (such as a large transformer) needing special permits and routing logistics.
SuperNIC
An AI-fabric network adapter built to spray packets across paths and reorder them in silicon; its rate is quoted per direction, so normalize before comparing it with bidirectional scale-up figures.
SuperPOD · DGX SuperPOD
NVIDIA's reference cluster design wiring many DGX systems into a validated, scalable AI supercomputer.
Switchgear
Assembly of breakers, switches and protection that controls and isolates electrical circuits; a long-lead procurement item.
SXM
NVIDIA's mezzanine socket form factor for GPUs, carrying full NVLink connectivity and higher power than PCIe cards, and the unit that server and price comparisons are quoted against.
Systolic array
A grid of processing elements that pumps data through in lockstep, the core structure of TPUs and many AI ASICs.
Tail latency · p99
The slowest few percent of responses (e.g. 99th percentile); in clusters one slow node can stall a whole job.
Take-or-pay
Contract obligating the buyer to pay for a minimum quantity of power or capacity whether or not it is used.
Tape-out
The milestone of finalizing a chip design and sending it to the foundry for fabrication.
TCO · Total Cost of Ownership
The full lifetime cost of capacity, including capex amortization, power, cooling, staff and maintenance.
TCP · Transmission Control Protocol
The ubiquitous reliable IP transport; it runs on any network without lossless tuning but costs latency, which is why hot GPU data paths use RDMA and reserve TCP for reach and simplicity.
TCS · Technology Cooling System
The secondary liquid loop that carries heat from IT equipment to the CDU; the facility loop rejects that heat onward to the outside world.
TDECQ · Transmitter Dispersion Eye Closure Quaternary
The dB measure of how far a transmitter's eye is already closed before the channel; margin erodes as case temperature rises, so qualify optics at field temperature rather than on a cool bench.
TDP · Thermal Design Power
A vendor-defined thermal/design rating used to size a product's power and cooling envelope. Its definition and test basis vary by product; TDP is not necessarily measured sustained electrical input, typical workload draw, or maximum/peak power.
TDX · Trust Domain Extensions
Intel's confidential-VM technology, enforced by a signed firmware module mediating hypervisor transitions; the CPU TEE choice that sets your GPU confidential-computing and attestation path.
TEE · Trusted Execution Environment
A hardware-isolated, encrypted region of a processor that protects code and data even from the host operator.
Tensor parallelism · TP
Splitting a single layer's math across multiple GPUs so they jointly compute one forward/backward pass.
TensorRT-LLM
NVIDIA's optimized library for compiling and serving large language models at low latency on its GPUs.
Test-time compute
Spending extra inference compute (e.g. chain-of-thought reasoning) to improve answers, shifting cost from training to serving.
TFLOPS · trillion floating-point operations per second
The throughput ladder for accelerators, rising through PFLOPS and EFLOPS; always read alongside the precision and sparsity assumptions, because the headline figure quotes the most favorable ones.
THD · Total harmonic distortion
The share of waveform distortion injected by non-linear loads; it is capped at the point of common coupling, which drives active-front-end, filter, and transformer decisions.
The cascade · Training-to-inference
The lifecycle where today's training hardware becomes tomorrow's inference fleet as newer chips arrive.
Thermal runaway
A self-reinforcing temperature rise, notably in batteries, that can lead to fire if not detected and contained.
TIA · Telecommunications Industry Association
The association behind the TIA-942 data-center standard, whose Rated scale is a separate certification from the Uptime Tiers buyers often assume it matches.
TIA-942
A telecom-industry standard for data-center infrastructure with its own Rated 1-4 reliability classification.
TiB · tebibyte
A binary-prefix capacity unit equal to 2^40 bytes, roughly ten percent larger than a decimal TB; mixing the two silently misstates storage and memory sizing.
Tier III
An Uptime classification meaning concurrently maintainable: any component can be serviced without taking IT load down.
Tier IV
The top Uptime classification meaning fault tolerant: the facility survives any single failure with no impact.
TIM · Thermal Interface Material
The paste or pad filling microscopic gaps between a chip and its heat spreader or cold plate to conduct heat.
Time-to-power · Speed-to-power
The elapsed time from contract to energized megawatts; the binding constraint and primary siting screen of the AI era.
TLC · Triple-Level Cell
Flash storing three bits per cell, the mainstream balance of cost, endurance and performance for SSDs.
Tokens-per-joule
Inference energy efficiency for a controlled run: output tokens divided by joules at a declared measurement boundary. Report model/version, workload, token accounting, precision/quality, context, batch/concurrency, latency, and utilization; it is not universal across setups.
Tokens-per-watt
Throughput per power under the same controlled conditions and boundary: (tokens/s)/W = tokens/J. “Tokens/W” is shorthand only when model/version, workload, token accounting, precision/quality, context, batch/concurrency, latency, utilization, and measurement boundary are stated.
Topology-aware scheduling
Placing a job's GPUs to respect network topology so its collectives run on high-bandwidth, low-latency links.
ToR · Top-of-Rack
The switch installed in each rack that aggregates that rack's server links before the leaf and spine layers; also shorthand for that whole layer of the topology.
TPOT · Time Per Output Token
The steady-state delay between successive generated tokens; the inter-token latency SLO governed by decode.
TPU · Tensor Processing Unit
Google's custom AI accelerator chip, built around a systolic array for matrix math and used across its cloud and models.
Trainium
Amazon's custom AI training accelerator, part of its bid to reduce dependence on merchant GPUs.
Transformer · Parts 2, 3, 4, 5, 6, 11, 12, 13, 14, 15, 16
Electrical machine that steps voltage between grid, distribution, and rack levels; among the longest-lead-time items in a build and a frequent schedule gate.
TrendForce
A market-research firm covering memory, packaging, and component supply, cited for high-bandwidth-memory and cooling-component pricing and availability.
Truck roll
Dispatching a technician on-site to fix something; minimizing truck rolls is a goal of remote and automated ops.
TSMC · Taiwan Semiconductor Manufacturing Company
The foundry that manufactures essentially all leading-edge AI logic and its advanced packaging, making its capacity the upstream gate on accelerator supply.
TSV · Through-Silicon Via
A vertical electrical channel drilled through a die to stack chips, the wiring that makes HBM and 3D stacking possible.
TTFT · Time To First Token
How long an inference request waits before the first output token appears; a key latency SLO set by prefill.
TUE · Total Usage Effectiveness
PUE multiplied by IT-side efficiency (ITUE); the true facility-to-transistor energy ratio.
Two-person rule
A control requiring two authorized people to act together for a sensitive operation, reducing insider risk.
Two-phase cooling
Cooling that absorbs heat by boiling a fluid and condensing it, exploiting latent heat for very high heat flux.
UALink
An open scale-up interconnect standard for up to 1,024 accelerators, the multi-vendor alternative to NVLink.
UCCM
The congestion-control mechanism in the Ultra Ethernet transport, paired with packet trimming so loss becomes a fast signal rather than a stall on a sprayed multipath fabric.
UCIe · Universal Chiplet Interconnect Express
An open standard for connecting chiplets from different vendors within one package.
UEC · Ultra Ethernet Consortium
The industry group defining AI-grade Ethernet transport (UET) with packet spray and modern congestion control.
UET · Ultra Ethernet Transport
The Ultra Ethernet Consortium's from-scratch AI transport, combining packet spray, NIC-side reordering, modern congestion control, and native RDMA as the open alternative to proprietary fabrics.
UL · Underwriters Laboratories
The safety-certification body whose listing makes equipment routinely acceptable to code officials and insurers; unlisted gear waits, however good its engineering case.
UPS · Uninterruptible Power Supply
System (battery or flywheel backed) that maintains clean power through grid disturbances and bridges to generators.
Uptime Tier
Uptime Institute's I-IV classification of facility resilience, from basic (I) to fault-tolerant 2N (IV).
UQD · Universal Quick Disconnect
A dripless connector letting liquid-cooled hardware be plugged and unplugged without spilling coolant or tools.
VAC · Volts alternating current
The unit naming conventional AC distribution down to server power supplies, the incumbent baseline that facility-level DC architectures are measured against.
VAR · Volt-ampere reactive
The unit of reactive power; equipment is rated in apparent power, so VAR handling rather than kilowatts alone sets generator, UPS, and transformer sizing.
VAST
A storage-platform vendor supplying all-flash file and object namespaces for AI clusters, and a common published source on checkpoint bandwidth requirements.
VBIOS · video BIOS
The GPU's own boot firmware; a stale or mismatched image produces gray failures such as link flaps and intermittent throttling that degrade a job without crashing the node.
VDC · Volts direct current
The unit naming rack and facility DC-bus architectures, whose higher levels delete conversion stages and copper at the cost of new protection and grounding engineering.
VFD · Variable-frequency drive
The drive that varies pump and fan speed; its slew limit bounds how fast cooling can answer a load step, and it is a harmonic source.
VLAN · Virtual Local Area Network
Layer-2 network segmentation; useful for zoning management and front-end traffic, but a configuration boundary rather than a cryptographic one, so never treat it as a hard security boundary.
vLLM
A popular open-source inference engine known for PagedAttention and high-throughput continuous batching.
VPC · Virtual Private Cloud
A VPC is a logically isolated tenant network; a DPU can enforce parts of its networking/security boundary at the host edge. Isolation still depends on DPU firmware and secure boot, control-plane authorization, configuration, shared-resource separation, updates, and adversarial testing—not hardware placement alone.
VPPA · Virtual Power Purchase Agreement
A financial contract for differences settled against a market price rather than delivering electrons, so generation shape and nodal basis can turn a hedge into a liability.
VRAM · video RAM
The memory attached to a GPU, used loosely as a synonym for its HBM; it is the ceiling that weights and the KV-cache compete for during inference.
VRLA · Valve-regulated lead-acid
The legacy sealed lead-acid UPS battery chemistry: cheap, heavy, and short-lived, but backed by a mature closed-loop recycling chain when the estate is decommissioned.
VRM · Voltage Regulator Module
Power-electronics stage that steps board voltage down to the low voltage a GPU or CPU core actually needs.
VXLAN · Virtual Extensible LAN
Overlay encapsulation that stretches isolated tenant networks across a shared routed fabric; a workhorse of multi-tenant isolation.
W45
ASHRAE's liquid-cooling classes keyed to maximum facility-water supply temperature, where a higher number means warmer water: less chiller capex and heat-reuse potential, less thermal headroom.
WACC · Weighted Average Cost of Capital
The blended cost of a project's debt and equity, used as the discount rate for valuing its cash flows.
WAN · Wide Area Network
The long-haul network between campuses; once a run outgrows any single powered site, WAN latency and bandwidth become a training-architecture constraint rather than an IT concern.
Warm-water loop
A liquid-cooling loop run at elevated temperature so heat can be rejected with free cooling and reused downstream.
Water-positive
A commitment to replenish more water than a facility consumes, a sustainability pledge by several operators.
WebDataset
A shard-based training data format that packs many small samples into large sequential archives, so the data loader reads streams instead of hammering the file system's metadata plane.
WEKA
A high-performance parallel file-system vendor whose platform is a common hot storage tier feeding GPU clusters over RDMA.
Wet-bulb temperature
The lowest temperature achievable by evaporation, setting the floor for how well evaporative cooling can perform.
White space
The conditioned data-hall area where IT racks sit, as opposed to support (gray) space for power and cooling gear.
WORM · Write Once Read Many
Storage that prevents data from being altered or deleted after writing, used for compliance and tamper resistance.
WUE · Water Usage Effectiveness
Liters of water consumed per kWh of IT energy; the water analog of PUE that evaporative cooling worsens.
XID
An NVIDIA GPU error code reported by the driver; specific XIDs flag memory, hardware or driver faults to triage.
XLA · Accelerated Linear Algebra
Google's ahead-of-time compiler for machine-learning graphs and the software gateway to TPUs; adopting it is a lock-in decision comparable to adopting CUDA or Neuron.
XPU
Generic term for a non-GPU AI accelerator such as a TPU, Trainium, Maia or MTIA; hyperscaler custom silicon.
Young/Daly
The formula setting the optimal checkpoint interval by balancing checkpoint cost against expected failure-rollback loss.
ZeRO · Zero Redundancy Optimizer
A technique partitioning optimizer state, gradients and parameters across GPUs to remove memory redundancy in training.
Zero trust
A security model that trusts no user or device by default and verifies every access request continuously.
Zero-touch provisioning · ZTP
Automatically configuring devices on first power-up with no manual setup, key to deploying at scale.
ZLD · Zero Liquid Discharge
A water system that recovers nearly all wastewater for reuse, leaving essentially no liquid discharge.