← Back to catalog

NVIDIA GB200 NVL72, an official NVIDIA product photo

NVIDIA GB200 NVL72

A full rack-scale system: 72 Blackwell GPUs plus 36 Grace CPUs, wired together with an NVLink switch fabric so the whole rack behaves like one giant computer with roughly 13.5 TB of shared GPU memory. This is genuine engineered clustering, not a pile of separate machines — see "What's a cluster?" for the difference.

In Stock (2) Ships from United States
Price
$3,000,000
This is the price of the hardware itself, in US dollars. It does not include electricity, cooling, networking gear, or setup labor.
GPU count
72
This is how many separate GPU chips are inside the system. Multiple GPUs can be wired together (see "What's a cluster?") so their memory and computing power combine to run bigger models.
GPU memory (per GPU)
192 GB
GPU memory (also called VRAM) is the fast memory built into the graphics card itself. The AI model has to fit entirely (or almost entirely) inside this memory to run at a decent speed. This is usually the single most important number when you're buying hardware to run a specific AI model — if the model doesn't fit, it either won't load, or it runs painfully slowly.
GPU memory (total)
13824 GB
GPU memory (also called VRAM) is the fast memory built into the graphics card itself. The AI model has to fit entirely (or almost entirely) inside this memory to run at a decent speed. This is usually the single most important number when you're buying hardware to run a specific AI model — if the model doesn't fit, it either won't load, or it runs painfully slowly.
CPU
36 NVIDIA Grace CPUs (Arm-based) built into the same rack — again, not the bottleneck for which models fit.
The CPU ("central processing unit") is the traffic-controller chip: it loads data and hands work to the GPUs. For running AI models, the CPU matters far less than GPU memory — a modest CPU is fine as long as the GPUs are sized correctly for the model.
System RAM
17280 GB
RAM is the computer's regular working memory (different from GPU memory). It holds data that isn't currently being crunched by the GPU. For running AI models, you mainly just need "enough that the CPU isn't waiting around" — it is not the number that decides whether a model fits.

What this power draw means

120000 W

Power draw is how much electricity the hardware pulls when it's working hard, measured in watts (W). It affects your electricity bill, and at large scale, whether your building's power and cooling can even support the equipment.

  • 100.0 average homes' worth of continuous power — To make a watt number feel real: we compare it to an average home, which draws roughly 1,200 watts around the clock (lights, fridge, HVAC, electronics, etc., averaged over a day). Dividing hardware wattage by 1,200 tells you how many average homes' worth of continuous power the hardware uses.
  • would drain a typical 90 kWh EV battery in about 0.8 hours — To make power draw feel real over time, we compare it to a typical electric car battery, which holds about 90 kWh (kilowatt-hours) of energy. That tells you roughly how many hours of running the hardware it would take to burn through one full EV battery's worth of energy.

Performance & fit

How this product actually measures up against the three big open models this store advises on — using the same memory math and build-cost numbers shown on each model's page, nothing new invented here.

  • Capable Can run Kimi K2.6 , though our advisor page recommends a cheaper option for the bare minimum build.
  • Capable Can run GLM-5.2 , though our advisor page recommends a cheaper option for the bare minimum build.
  • Capable Can run DeepSeek V4 Flash , though our advisor page recommends a cheaper option for the bare minimum build.

Where these numbers come from

GPU count and total memory (13.5 TB HBM3e, i.e. 192 GB per GPU) from NVIDIA's official GB200 NVL72 product page. Power draw: NVIDIA's nominal spec is 120 kW per rack, though real-world deployments have been reported running 130–132 kW under full load — we show the 120 kW nominal figure but flag it as an estimate. System RAM (36 Grace CPUs) is computed from the Grace CPU's published 480 GB LPDDR5X spec × 36, not a single number NVIDIA states for the whole rack.

Shipping information

Ships from United States. Rack- and server-scale systems are built and shipped to order — freight, power, and cooling requirements vary a lot by site, so contact our team for a shipping quote and delivery timeline specific to your installation.

Return policy

Custom-built systems at this scale aren't eligible for a standard return window. If a unit arrives damaged or doesn't match the agreed specification, contact us within 15 days of delivery and we'll work with you to resolve it.

Warranty

Covered by NVIDIA's enterprise hardware warranty and support options for datacenter and rack-scale systems, with multi-year extended support contracts available. Exact terms are confirmed as part of your quote.

Request this product

Tell us your name and email and we'll follow up about this exact package — no payment, no account needed.

You might also like

NVIDIA DGX B200, an official NVIDIA product photo
Server

NVIDIA DGX B200

A complete 8-GPU AI server — the standard building block for serious AI infrastructure.

$400,000
1440 GB GPU memory (8× GPU)
10200 W
NVIDIA GB300 NVL72, an official NVIDIA product photo
Rack

NVIDIA GB300 NVL72

NVIDIA's current flagship rack: 72 Blackwell Ultra GPUs, built for the largest models at scale.

$4,000,000
20736 GB GPU memory (72× GPU)
135000 W
NVIDIA H200 (SXM5), an official NVIDIA product photo
Datacenter Gpu

NVIDIA H200 (SXM5)

A single datacenter-grade AI GPU — the kind that gets wired into 8-GPU servers.

$35,000
141 GB GPU memory
700 W
RTX PRO 6000 Blackwell Workstation Edition, an official NVIDIA product photo
Workstation Gpu

RTX PRO 6000 Blackwell Workstation Edition

A professional workstation card with three times the memory of a gaming card.

$8,565
96 GB GPU memory
600 W