Air-cooled GPU rows still dominate many AI halls. They also hit a hard ceiling once you try to put an entire GB200 NVL72 class rack on the floor: 72 Blackwell GPUs and 36 Grace CPUs in one liquid-cooled NVLink domain, with NVIDIA stating about 120 kW of rack power on DGX GB200 hardware docs and OEM builds listing 125–135 kW operating power on the Supermicro NVL72 datasheet.
This full guide walks through what GB200 NVL72 actually is, how the liquid-cooled copper scale-up fabric works, what facility power and CDUs have to deliver, where HGX air-cooled still wins, and which IDT guides cover the adjacent bottlenecks (800 VDC, liquid cooling, NVLink vs InfiniBand vs UALink).
Key Takeaways
- GB200 NVL72 is a liquid-cooled rack: 36 Grace + 72 Blackwell in one 72-GPU NVLink domain.
- NVIDIA quotes 130 TB/s domain NVLink bandwidth and 1.8 TB/s per GPU.
- Plan for roughly 120 kW to 125–135 kW continuous IT load per rack, not air-row density.
- Cooling is mandatory: cold plates, manifolds, CDU (OEM 250 kW in-rack options), leak detection.
- Honest limits: facility power, coolant loops, CoWoS/HBM lead times, and CUDA-stack lock-in still decide who can deploy.
What GB200 NVL72 actually is
Per NVIDIA's GB200 NVL72 product page, the rack connects 36 NVIDIA Grace CPUs and 72 Blackwell GPUs in a liquid-cooled, rack-scale design. The headline architectural bet is a 72-GPU NVLink domain that behaves like one massive accelerator for trillion-parameter LLM inference and training.
The building block is the GB200 Grace Blackwell Superchip: 1 Grace CPU plus 2 Blackwell GPUs, linked by NVLink-C2C at 900 GB/s bidirectional according to NVIDIA's March 2024 technical blog. A compute tray packs two Superchips (2 Grace + 4 Blackwell), and an NVL72 rack stacks 18 of those trays with 9 NVLink switch trays, as detailed in the DGX GB rack hardware guide.
NVIDIA also ships smaller NVL36 configurations and two-rack NVL72 layouts. Operators should read the SKU sheet before assuming every "GB200" quote is a single-rack 72-GPU domain.
Specs that matter for buyers
The numbers below come from NVIDIA's published NVL72 table and OEM rack datasheets. Sparse versus dense FLOPS footnotes matter; treat marketing multipliers as workload-specific, not universal.
| Metric | GB200 NVL72 (NVIDIA / OEM) | Why it matters |
|---|---|---|
| GPUs / CPUs | 72 Blackwell + 36 Grace | Defines rack-scale NVLink domain size |
| NVLink domain bandwidth | 130 TB/s | All-to-all GPU fabric inside the rack |
| Per-GPU NVLink | 1.8 TB/s bidirectional | Fifth-generation NVLink scale-up |
| GPU memory | 13.4 TB HBM3E at 576 TB/s | Unified fast memory pool framing |
| CPU memory | 17 TB LPDDR5X at 14 TB/s | Grace host pool for CPU-side work |
| NVFP4 (sparse | dense) | 1,440 | 720 PFLOPS | Headline inference precision path |
| Rack power (NVIDIA DGX) | approx. 120 kW | Facility feed and bus-bar design |
| Operating power (Supermicro) | 125–135 kW; shelves 132 kW | Real continuous IT planning band |
Source: NVIDIA GB200 NVL72 specs; DGX GB200 hardware; Supermicro SuperCluster GB200 NVL72 datasheet.
For memory packaging context, HBM stacks and CoWoS allocation still gate how many Blackwell GPUs the industry can ship. See Inside Deep Tech's HBM full guide and TSMC CoWoS packaging guide.
Why the 72-GPU NVLink domain is the product
Before NVL72, the practical NVLink domain ceiling on HGX-class boards was eight GPUs. NVIDIA's multi-node NVLink tuning overview states that fifth-generation NVLink raises per-GPU bandwidth to 1.8 TB/s and expands the domain to 72 Blackwell GPUs inside one NVL72 rack.
Inside the rack, nine NVLink switch trays plus a copper cable cartridge stitch the GPUs together. NVIDIA's product page puts aggregate domain bandwidth at 130 TB/s. That is scale-up fabric, not a substitute for scale-out networking. Cluster-level training and multi-rack inference still ride Quantum InfiniBand or Spectrum-X Ethernet, a split covered in the NVLink, InfiniBand, and UALink guide.
Here's why that matters. MoE and trillion-parameter inference burn tokens when activations bounce over PCIe or Ethernet hops. Collapsing 72 GPUs into one NVLink domain is NVIDIA's answer when copper board-level scale-up runs out of sockets.
Performance claims, with NVIDIA's own caveats
NVIDIA markets NVL72 with three headline multipliers versus H100-class baselines on the product page: 30x real-time trillion-parameter LLM inference, 4x LLM training, and 25x energy efficiency versus H100 air-cooled infrastructure at the same power.
The technical blog spells out inference assumptions: token-to-token latency around 50 ms, first-token latency around 5,000 ms, input sequence 32,768, output 1,024, comparing air-cooled HGX H100 over InfiniBand against liquid-cooled GB200 Superchips in NVL72. Those are workload-bound projections, not a guarantee for every model or batch size.
Operators should re-benchmark their own MoE and dense checkpoints. FP4 / NVFP4 paths, KV-cache behavior, and NCCL topology sensitivity all change the shape of the curve.
Liquid cooling is not optional on NVL72
NVL72 is a direct-to-chip liquid-cooled design. The DGX hardware guide describes manifolds feeding cold plates on CPUs and GPUs, with NICs and storage still air-cooled by tray fans. Leak detection is treated as a first-class reliability feature, not a nice-to-have.
OEM packaging makes the facility ask concrete. Supermicro's NVL72 datasheet lists an in-rack 250 kW CDU with redundant PSU and dual hot-swap pumps, a 1.3 MW in-row CDU option, and 180 kW / 240 kW liquid-to-air solutions when facility water is unavailable. GIGABYTE's GIGAPOD NVL72 page additionally points integrators at ASHRAE liquid-cooling water-quality guidance for the facility water system.
For the broader CDU, cold-plate, and immersion landscape, use Inside Deep Tech's data center liquid cooling full guide. NVL72 is the canonical reason many 2026 AI halls stop pretending rear-door heat exchangers alone will save a 120 kW+ row.
Power path: bus bars, shelves, and why 48 V starts to hurt
NVIDIA documents a bus-bar distribution path: power shelves convert AC to roughly 50–51 V DC and feed the trays, with about 120 kW rack draw, eight shelves, six 5.5 kW PSUs per shelf, and 33 kW per shelf under N+N redundancy (DGX GB200 hardware). Supermicro publishes matching shelf math at 132 kW total with 125–135 kW operating power (datasheet).
That density is exactly why Inside Deep Tech tracks 800 VDC power distribution and SiC/GaN power electronics for AI racks. Copper cross-section and conversion stages get painful as continuous IT load climbs past 100 kW per rack. NVL72 does not invent the power problem; it makes the problem impossible to ignore.
If your campus still waits on utility interconnection for multi-hundred-megawatt AI blocks, pair this guide with the behind-the-meter generation full guide. A liquid-cooled GPU rack without electrons is just expensive plumbing.
When HGX air-cooled (and smaller NVL domains) still win
NVL72 is the right answer when the model and the facility both demand rack-scale NVLink. It is the wrong answer when any of these are true:
- The hall cannot deliver roughly 120 kW continuous per rack with a commissioned coolant loop.
- The workload fits well in 8-GPU HGX domains and does not need a 72-GPU NVLink island.
- Ops teams lack leak-detection, CDU redundancy, and copper-cartridge field procedures.
- Procurement needs faster air-cooled B200 / HGX capacity while NVL72 lead times slip.
Supermicro's October 2024 liquid-cooled SuperCluster announcement itself framed HGX B200 liquid- and air-cooled systems as parallel SKUs. That portfolio split is the tell: NVL72 is not "the only Blackwell," it is the densest NVLink product line.
Software and lock-in: CUDA, NCCL, and the open-fabric debate
NVL72 assumes the NVIDIA software stack: CUDA, NCCL, Magnum IO, and increasingly factory-ops tooling such as Mission Control on the product page. That is a strength when your teams already ship on CUDA. It is a constraint if you are evaluating multi-vendor scale-up fabrics.
UALink and Ultra Ethernet exist precisely because buyers want an escape hatch from proprietary scale-up. Inside Deep Tech covers that tension in the interconnect guide and the UALink full guide. For 2026 deployments, most production NVL72 clusters will still be CUDA-first. Plan staffing and ISV certification accordingly.
Supply chain reality: HBM, CoWoS, and lead times
Even a perfect facility plan fails if Blackwell dies and HBM stacks are allocation-constrained. CoWoS packaging capacity and HBM3E supply remain industry bottlenecks; see TSMC CoWoS and HBM. Quote NVL72 as a system (trays, switches, CDU, networking), not a GPU unit price.
What a real NVL72 deployment checklist looks like
Field teams treat NVL72 as a mini campus project, not a server rack. A practical sequence mirrors OEM and integrator practice:
- Confirm floor loading, whip ampacity, and UPS/generator headroom for continuous 120 kW-class draws plus CDU pumps.
- Flush, fill, purge, and pressure-test the liquid loop before power-on; commission leak detection and pump failover.
- Seat every copper NVLink cartridge and verify NVLink Switch tray health before NCCL burn-in.
- Bring up management TOR, BMC paths, and BlueField / ConnectX links documented in the DGX hardware guide.
- Run thermal soak under sustained all-reduce and inference loads; watch supply/return deltas and PSU redundancy pull-tests.
Skip any of those steps and you inherit intermittent thermal trips that look like "software bugs" for weeks. Liquid-cooled AI is an operations discipline.
Density math: why air rows run out of copper and CFM
A conventional air-cooled GPU row often tops out well below 40–60 kW per rack once you account for hot aisle containment, CRAH capacity, and cable bulk. NVL72 asks for roughly 125–135 kW of IT load in one footprint. That is not a 2x jump. It is a different building type.
Copper bus bar cross-section, breaker coordination, and conversion losses all scale with current. That is the engineering reason Inside Deep Tech keeps returning to 800 VDC distribution for future AI halls: at 100 kW+ continuous, low-voltage DC starts eating copper and efficiency. NVL72 deployments in 2026 still mostly land on today's 415/480 VAC feeds into power shelves, but the trajectory is clear.
Cooling air volume also collapses as a strategy. Moving tens of kilowatts through heatsinks and fans fights physics. Direct-to-chip liquid rejects heat at the die, which is why NVIDIA pairs NVL72 with cold plates instead of hoping rear-door exchangers catch an entire 72-GPU domain.
Where NVL72 sits versus GB300 and the roadmap
NVIDIA's product family already points past GB200 toward GB300 NVL72 and Vera Rubin class racks on the same product hub. Buyers should assume NVL72 is a multi-year platform, not a one-quarter SKU, but they should also assume firmware, networking, and cooling revisions will keep moving.
The durable lesson is architectural: rack-scale NVLink domains plus liquid cooling become the default for frontier training and real-time trillion-parameter inference. HGX remains the workhorse for many enterprise fine-tunes and mid-size serving fleets. Portfolio planning should hold both lanes.
Operator risks that do not show up in FLOPS tables
Three risks dominate post-PO regret:
- Facility water chemistry and CDU fouling (ASHRAE guidance called out by GIGABYTE and peers).
- Partial-rack failures that strand an entire 72-GPU NVLink island if switch trays or cartridges are single points of delay.
- Team skill mix: liquid-loop technicians, NVLink fabric debugging, and CUDA performance engineering rarely live in one hire.
Budget for spare cold plates, QD fittings, and trained on-call coverage the same way you budget spare GPUs. The mean time to recover a liquid leak event is an ops metric, not a vendor slide.
How to read OEM datasheets without getting played
When comparing Supermicro, Dell, GIGABYTE, and other MGX partners, normalize four fields: continuous operating power, CDU topology (in-rack vs in-row vs liquid-to-air), networking SKUs (Quantum vs Spectrum-X), and what is included versus "customer-furnished." A 250 kW CDU rating is capacity headroom, not a promise your facility water loop can reject that heat on a summer day.
Also separate NVIDIA's projected 30x / 4x / 25x multipliers from your acceptance tests. Contract language should reference your model, sequence lengths, and latency SLOs, not a blog chart.
Inside Deep Tech's take
Inside Deep Tech's take: GB200 NVL72 is the clearest proof that AI infrastructure has left the air-cooled server era. The interesting question is no longer "how many GPUs fit in 8U." It is whether your power, water, and copper plans can host a 120 kW+ NVLink island without turning the rest of the campus into a science project.
Buy NVL72 when the model needs the domain and the hall is ready. Buy HGX (or wait) when either condition fails. Do not paper over a facility gap with a denser SKU.
FAQ
What is NVIDIA GB200 NVL72?
GB200 NVL72 is NVIDIA's liquid-cooled rack-scale system that connects 36 Grace CPUs and 72 Blackwell GPUs into one 72-GPU NVLink domain for large-model AI training and inference.
How much power does a GB200 NVL72 rack use?
NVIDIA's DGX GB200 hardware guide cites roughly 120 kW rack power. Supermicro's datasheet lists 125–135 kW operating power with 132 kW of power-shelf capacity.
Does GB200 NVL72 require liquid cooling?
Yes. The compute path uses direct-to-chip cold plates and rack manifolds. OEM options include in-rack 250 kW CDUs and liquid-to-air units when facility water is unavailable. See also Inside Deep Tech's liquid cooling guide.
How is NVLink bandwidth quoted on NVL72?
How does NVL72 differ from HGX B200?
HGX B200 is typically an 8-GPU board-level NVLink domain suited to air- or liquid-cooled servers. NVL72 expands to a 72-GPU rack-scale domain with Grace CPUs, copper cartridges, and mandatory liquid cooling for the dense compute path.
What memory capacity does the rack advertise?
When should a team still buy air-cooled HGX instead?
When the facility cannot host ~120 kW liquid-cooled racks, when the workload fits 8-GPU domains, or when lead times and ops maturity favor simpler air-cooled deployments.
How does NVL72 connect to campus networking?
Scale-up stays on NVLink inside the rack. Scale-out uses NVIDIA Quantum InfiniBand or Spectrum-X Ethernet, plus ConnectX / BlueField adapters documented in the DGX hardware guide. See the interconnect full guide for the fabric split.
What to do Monday
Map every proposed NVL72 row to three numbers: continuous kW per rack, CDU capacity and redundancy, and NVLink-versus-scale-out topology. If any of the three is hand-wavy, pause procurement and fix the facility package first.
Then cross-check memory and packaging risk (HBM, CoWoS) and the power conversion path (800 VDC, SiC/GaN). NVL72 only pays off when the whole stack arrives together.



