> ## Content Index
> Fetch the complete content index at: https://www.insidedeeptech.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# NVIDIA Vera Rubin: A Full Guide to the Rubin GPU, Vera CPU, and NVL72 Racks After Blackwell
- URL: https://www.insidedeeptech.com/nvidia-vera-rubin-nvl72-full-guide/
- Published: 2026-10-05T06:11:50.000Z
- Updated: 2026-10-05T06:11:50.000Z
- Description: NVIDIA Vera Rubin full guide (Oct 2026): Rubin GPU, Vera CPU, NVL72 racks now shipping, HBM4 suppliers, NVLink 6, Groq 3 LPX, the NVL144 naming change, Rubin Ultra and Kyber, and honest limits.
- Author: Austin Heaton
- Tags: AI, Hardware, Semiconductors, Deep Tech

Blackwell taught the industry to buy AI compute by the rack. [NVIDIA Vera Rubin](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com) is the first generation where the GPU, the CPU, the memory, the rack and the network were all redesigned at once, and as of October 5, 2026 it's no longer a roadmap slide: NVIDIA told investors it [commenced production shipments of Vera Rubin in August 2026](https://s201.q4cdn.com/141608511/files/content%5Ffiles/TRANSCRIPT%5F-NVIDIA-Corp-NVDA-US-Q2-2027-Earnings-Call-26-August-2026-5%5F00-PM-ET.pdf?ref=insidedeeptech.com).

So the useful question isn't what Rubin is. It's what Rubin changes. The short answer is memory and plumbing: per NVIDIA's [Rubin GPU architecture deep dive](https://developer.nvidia.com/blog/inside-nvidia-rubin-gpu-architecture-powering-the-era-of-agentic-ai/?ref=insidedeeptech.com), each GPU carries up to **288 GB** of HBM4 at up to **22 TB/s**, a **2.8x** jump over Blackwell, and the [Vera Rubin NVL72 rack](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com) ties **72** of those GPUs to **36** Vera CPUs over NVLink 6.

This full guide covers the Rubin GPU, the Vera CPU, the NVL72 rack (and why it used to be called NVL144), HBM4, networking, the Rubin Ultra and Kyber roadmap, and the honest limits. If you're still deploying the previous generation, start with Inside Deep Tech's [GB200 NVL72 full guide](https://www.insidedeeptech.com/gb200-nvl72-nvidia-ai-rack-full-guide/).

## Key Takeaways

- Vera Rubin NVL72 pairs [72 Rubin GPUs with 36 Vera CPUs](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com); [production shipments began in August 2026](https://s201.q4cdn.com/141608511/files/content%5Ffiles/TRANSCRIPT%5F-NVIDIA-Corp-NVDA-US-Q2-2027-Earnings-Call-26-August-2026-5%5F00-PM-ET.pdf?ref=insidedeeptech.com).
- Each Rubin GPU lists [288 GB HBM4 and 50 PFLOPS NVFP4](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com) (sparse); bandwidth is quoted at [19.2 TB/s](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com) or [up to 22 TB/s](https://developer.nvidia.com/blog/inside-nvidia-rubin-gpu-architecture-powering-the-era-of-agentic-ai/?ref=insidedeeptech.com).
- NVLink 6 doubles per-GPU scale-up to [3.6 TB/s](https://developer.nvidia.com/blog/inside-the-nvidia-rubin-platform-six-new-chips-one-ai-supercomputer/?ref=insidedeeptech.com); NVIDIA's spec table lists [216 TB/s per rack](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com), its POD blog [260 TB/s](https://developer.nvidia.com/blog/nvidia-vera-rubin-pod-seven-chips-five-rack-scale-systems-one-ai-supercomputer/?ref=insidedeeptech.com).
- "NVL144" and "NVL72" name the same 36-CPU rack class: [2025 materials counted 144 GPUs](https://nvidianews.nvidia.com/news/nvidia-unveils-rubin-cpx-a-new-class-of-gpu-designed-for-massive-context-inference?ref=insidedeeptech.com), [2026 materials count dual-die packages](https://developer.nvidia.com/blog/inside-nvidia-rubin-gpu-architecture-powering-the-era-of-agentic-ai/?ref=insidedeeptech.com).
- Honest limits: [supply stays a bottleneck through fiscal 2028](https://s201.q4cdn.com/141608511/files/content%5Ffiles/TRANSCRIPT%5F-NVIDIA-Corp-NVDA-US-Q2-2027-Earnings-Call-26-August-2026-5%5F00-PM-ET.pdf?ref=insidedeeptech.com), no public rack power figure, and Rubin CPX's status is unconfirmed.

⚠️

NVIDIA Vera Rubin is shipping, but not every number on NVIDIA's own pages agrees. Plan capacity on the conservative spec-table figures and treat keynote multipliers as vendor projections until your own benchmarks land.

## What NVIDIA Vera Rubin actually is in October 2026

Vera Rubin is a platform name, not a chip. NVIDIA's [GTC 2026 announcement](https://nvidianews.nvidia.com/news/nvidia-vera-rubin-platform?ref=insidedeeptech.com) on March 16, 2026 defined it as seven chips in full production: the Vera CPU, Rubin GPU, NVLink 6 Switch, ConnectX-9 SuperNIC, BlueField-4 DPU, Spectrum-6 Ethernet switch, and the newly integrated Groq 3 LPU.

Those chips ship in five rack types, per the same [GTC 2026 release](https://nvidianews.nvidia.com/news/nvidia-vera-rubin-platform?ref=insidedeeptech.com): Vera Rubin NVL72 GPU racks, Vera CPU racks, Groq 3 LPX inference racks, BlueField-4 STX storage racks, and Spectrum-6 SPX Ethernet racks. NVIDIA's [Vera Rubin POD blog](https://developer.nvidia.com/blog/nvidia-vera-rubin-pod-seven-chips-five-rack-scale-systems-one-ai-supercomputer/?ref=insidedeeptech.com) describes the reference POD as **40** racks with **1,152** Rubin GPUs and **60** exaflops.

NVIDIA's [CES 2026 launch release](https://nvidianews.nvidia.com/news/rubin-platform-ai-supercomputer?ref=insidedeeptech.com) also introduced the HGX Rubin NVL8 system for eight-GPU servers. Vera is the CPU. Rubin is the GPU. The NVL number is the size of the NVLink domain.

### Timeline: from roadmap slide to production shipments

| Date         | Milestone                                                                            | Source                                                                                                                                                                                                                                                                                                          |
| ------------ | ------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Sep 9, 2025  | Rubin CPX and the Vera Rubin NVL144 CPX rack announced; CPX targeted for end of 2026 | [NVIDIA Newsroom](https://nvidianews.nvidia.com/news/nvidia-unveils-rubin-cpx-a-new-class-of-gpu-designed-for-massive-context-inference?ref=insidedeeptech.com)                                                                                                                                                 |
| Sep 12, 2025 | SK hynix completes HBM4 development and readies mass production                      | [SK hynix](https://news.skhynix.com/en/sk-hynix-completes-worlds-first-hbm4-development-and-readies-mass-production/?ref=insidedeeptech.com)                                                                                                                                                                    |
| Jan 5, 2026  | Rubin platform launched at CES with six chips                                        | [NVIDIA Newsroom](https://nvidianews.nvidia.com/news/rubin-platform-ai-supercomputer?ref=insidedeeptech.com)                                                                                                                                                                                                    |
| Feb 12, 2026 | Samsung begins HBM4 mass production and ships commercial parts                       | [Samsung](https://news.samsungsemiconductor.com/global/samsung-ships-industry-first-commercial-hbm4-with-ultimate-performance-for-ai-computing/?ref=insidedeeptech.com)                                                                                                                                         |
| Mar 16, 2026 | GTC: seven chips in full production; Micron HBM4 in high-volume production           | [NVIDIA](https://nvidianews.nvidia.com/news/nvidia-vera-rubin-platform?ref=insidedeeptech.com), [Micron](https://investors.micron.com/news/press-release/2026/Micron-in-High-Volume-Production-of-HBM4-Designed-for-NVIDIA-Vera-Rubin-PCIe-Gen6-SSD-and-SOCAMM2-03-16-2026/default.aspx?ref=insidedeeptech.com) |
| May 31, 2026 | Full production ramp; shipments set to begin "this fall"                             | [NVIDIA Newsroom](https://nvidianews.nvidia.com/news/vera-rubin-full-production-agentic-ai-factory?ref=insidedeeptech.com)                                                                                                                                                                                      |
| Aug 2026     | Production shipments of Vera Rubin commence                                          | [Q2 FY27 call](https://s201.q4cdn.com/141608511/files/content%5Ffiles/TRANSCRIPT%5F-NVIDIA-Corp-NVDA-US-Q2-2027-Earnings-Call-26-August-2026-5%5F00-PM-ET.pdf?ref=insidedeeptech.com)                                                                                                                           |

Source: [NVIDIA Newsroom](https://nvidianews.nvidia.com/news/nvidia-vera-rubin-platform?ref=insidedeeptech.com), [NVIDIA Q2 fiscal 2027 earnings call transcript](https://s201.q4cdn.com/141608511/files/content%5Ffiles/TRANSCRIPT%5F-NVIDIA-Corp-NVDA-US-Q2-2027-Earnings-Call-26-August-2026-5%5F00-PM-ET.pdf?ref=insidedeeptech.com), and the supplier releases linked in the table.

Read those dates carefully. "Full production" in March meant chips. [Production shipments of systems](https://s201.q4cdn.com/141608511/files/content%5Ffiles/TRANSCRIPT%5F-NVIDIA-Corp-NVDA-US-Q2-2027-Earnings-Call-26-August-2026-5%5F00-PM-ET.pdf?ref=insidedeeptech.com) began in August, and NVIDIA's CFO guided that Vera Rubin would be about **20%** of data center revenue in the fiscal third quarter, per the [same call](https://s201.q4cdn.com/141608511/files/content%5Ffiles/TRANSCRIPT%5F-NVIDIA-Corp-NVDA-US-Q2-2027-Earnings-Call-26-August-2026-5%5F00-PM-ET.pdf?ref=insidedeeptech.com).

## The Rubin GPU: two reticle dies, HBM4, and a new Transformer Engine

Per NVIDIA's [Rubin GPU architecture blog](https://developer.nvidia.com/blog/inside-nvidia-rubin-gpu-architecture-powering-the-era-of-agentic-ai/?ref=insidedeeptech.com), published July 21, 2026, Rubin is two reticle-limited compute dies joined on one package by the NVIDIA High-Bandwidth Interface (NV-HBI). The package carries **336 billion** transistors, **224** streaming multiprocessors and **896** Tensor Cores.

The headline figure is up to **50 petaflops** of NVFP4 inference per GPU, which NVIDIA's [product spec table](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com) footnotes as a sparse number. The dense figures are the ones to plan around: **35 PFLOPS** NVFP4 training, **17.5 PFLOPS** FP8/FP6 and **4 PFLOPS** FP16/BF16 per GPU.

The architecture changes matter more than the peak. The [same deep dive](https://developer.nvidia.com/blog/inside-nvidia-rubin-gpu-architecture-powering-the-era-of-agentic-ai/?ref=insidedeeptech.com) lists:

- Tensor Cores that process twice the data along the K dimension per clock ([NVIDIA](https://developer.nvidia.com/blog/inside-nvidia-rubin-gpu-architecture-powering-the-era-of-agentic-ai/?ref=insidedeeptech.com)).
- A 3-bit lookup-table weight format that NVIDIA says can retain up to MXFP8 accuracy.
- 2:4 activation sparsity for attention, plus [2x FP32 and 4x BF16/FP16](https://developer.nvidia.com/blog/inside-nvidia-rubin-gpu-architecture-powering-the-era-of-agentic-ai/?ref=insidedeeptech.com) exponential throughput versus Blackwell for softmax.
- Counted writes that cut synchronization overhead on device-initiated NVLink transfers.

Per-GPU links, from the [Rubin architecture blog](https://developer.nvidia.com/blog/inside-nvidia-rubin-gpu-architecture-powering-the-era-of-agentic-ai/?ref=insidedeeptech.com): NVLink 6 at **3,600 GB/s**, NVLink-C2C to the Vera CPU at **1,800 GB/s**, and x16 PCIe Gen 6 at up to **256 GB/s**.

### Rubin GPU, Vera Rubin Superchip, and NVL72 rack specs

| Metric                               | Rubin GPU                                                                                      | Vera Rubin Superchip          | Vera Rubin NVL72                                                                                                  |
| ------------------------------------ | ---------------------------------------------------------------------------------------------- | ----------------------------- | ----------------------------------------------------------------------------------------------------------------- |
| Configuration                        | 1 Rubin GPU                                                                                    | 2 Rubin GPUs + 1 Vera CPU     | [72 Rubin GPUs + 36 Vera CPUs](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com) |
| NVFP4 inference (sparse)             | [50 PFLOPS](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com) | 100 PFLOPS                    | [3,600 PFLOPS](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com)                 |
| NVFP4 training (dense)               | 35 PFLOPS                                                                                      | 70 PFLOPS                     | [2,520 PFLOPS](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com)                 |
| FP8/FP6 training (dense)             | 17.5 PFLOPS                                                                                    | 35 PFLOPS                     | [1,260 PFLOPS](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com)                 |
| FP16/BF16 (dense)                    | 4 PFLOPS                                                                                       | 8 PFLOPS                      | 288 PFLOPS                                                                                                        |
| FP64                                 | 33 TFLOPS                                                                                      | 67 TFLOPS                     | [2,400 TFLOPS](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com)                 |
| HBM4 capacity / bandwidth            | 288 GB / 19.2 TB/s                                                                             | 576 GB / 38.5 TB/s            | [20.7 TB / 1,400 TB/s](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com)         |
| NVLink bandwidth                     | 3 TB/s                                                                                         | 6 TB/s                        | 216 TB/s                                                                                                          |
| CPU memory                           | n/a                                                                                            | Up to 1.5 TB LPDDR5X          | Up to 54 TB LPDDR5X                                                                                               |
| CPU cores                            | n/a                                                                                            | 88 Olympus cores, 176 threads | [3,168 Olympus cores](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com)          |
| Scale-out networking (bidirectional) | 0.45 TB/s                                                                                      | 0.9 TB/s                      | 32.4 TB/s                                                                                                         |

Source: [NVIDIA Vera Rubin NVL72 product page](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com), spec table (sparse and dense footnotes as published), retrieved October 5, 2026.

Notice the per-GPU NVLink figure. The spec table says **3 TB/s**, while NVIDIA's [platform deep dive](https://developer.nvidia.com/blog/inside-the-nvidia-rubin-platform-six-new-chips-one-ai-supercomputer/?ref=insidedeeptech.com) and [CES release](https://nvidianews.nvidia.com/news/rubin-platform-ai-supercomputer?ref=insidedeeptech.com) say **3.6 TB/s**. The conflicting-numbers section below explains how to handle that.

## The Vera CPU: 88 Olympus cores built for agent sandboxes

Vera succeeds Grace as NVIDIA's data center CPU. NVIDIA's [Vera CPU architecture blog](https://developer.nvidia.com/blog/inside-nvidia-vera-cpu-olympus-cores-built-for-maximum-single-threaded-performance-in-agentic-ai/?ref=insidedeeptech.com) describes **88** custom Olympus cores with full Arm compatibility, Spatial Multithreading that partitions core resources between two hardware threads, a **164 MB** unified L3 cache, and a Scalable Coherency Fabric with up to **3.4 TB/s** of bisection bandwidth.

Memory is SOCAMM2 LPDDR5X with up to **1.2 TB/s** of aggregate bandwidth in field-replaceable modules, per the [same blog](https://developer.nvidia.com/blog/inside-nvidia-vera-cpu-olympus-cores-built-for-maximum-single-threaded-performance-in-agentic-ai/?ref=insidedeeptech.com). NVIDIA's [spec table](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com) lists up to **1.5 TB** of LPDDR5X per Vera CPU, and Micron says its SOCAMM2 enables [up to 2TB and 1.2 TB/s per CPU](https://investors.micron.com/news/press-release/2026/Micron-in-High-Volume-Production-of-HBM4-Designed-for-NVIDIA-Vera-Rubin-PCIe-Gen6-SSD-and-SOCAMM2-03-16-2026/default.aspx?ref=insidedeeptech.com).

Why does a CPU matter this much? Agentic workloads push tool calls, code execution and reinforcement learning environments onto CPU cores. The [GTC 2026 release](https://nvidianews.nvidia.com/news/nvidia-vera-rubin-platform?ref=insidedeeptech.com) introduced a standalone Vera CPU rack with **256** liquid-cooled Vera CPUs, and NVIDIA's [POD blog](https://developer.nvidia.com/blog/nvidia-vera-rubin-pod-seven-chips-five-rack-scale-systems-one-ai-supercomputer/?ref=insidedeeptech.com) says one rack can sustain over **22,500** concurrent RL or agent sandbox environments.

Inside Deep Tech's take: the Vera CPU rack is this generation's sleeper product. It turns NVIDIA's CPU from a GPU companion into a product line aimed straight at x86 sandbox fleets.

## Inside the Vera Rubin NVL72 rack

The rack keeps the NVL72 footprint but rebuilds almost everything inside it. Per NVIDIA's [Vera Rubin POD blog](https://developer.nvidia.com/blog/nvidia-vera-rubin-pod-seven-chips-five-rack-scale-systems-one-ai-supercomputer/?ref=insidedeeptech.com), it holds **18** compute trays and **9** NVLink switch trays, about **1.3 million** components, and weighs roughly **4,000** lbs. Each compute tray carries two Vera Rubin superchips.

Trays are cable-free, hose-free and fanless. NVIDIA says tray assembly drops from nearly two hours to about five minutes, and the NVLink spine at the back uses four cable cartridges holding **5,000** copper cables ([POD blog](https://developer.nvidia.com/blog/nvidia-vera-rubin-pod-seven-chips-five-rack-scale-systems-one-ai-supercomputer/?ref=insidedeeptech.com)).

Power and cooling changed too. The same [POD blog](https://developer.nvidia.com/blog/nvidia-vera-rubin-pod-seven-chips-five-rack-scale-systems-one-ai-supercomputer/?ref=insidedeeptech.com) lists a **45°C** warm-water inlet, liquid-cooled busbars rated up to **5,000 A**, and Intelligent Power Smoothing with **6x** more rack-level energy storage (**400 J** per GPU) that cuts peak current demand by up to **25%**.

That 45°C figure matters for facilities. NVIDIA says a [45°C inlet](https://developer.nvidia.com/blog/nvidia-vera-rubin-pod-seven-chips-five-rack-scale-systems-one-ai-supercomputer/?ref=insidedeeptech.com) lets many climates use dry coolers instead of chillers, freeing enough power for up to **10%** more NVL72 racks in the same budget. For cold plates and CDUs, the Blackwell-era rules in Inside Deep Tech's [GB200 NVL72 guide](https://www.insidedeeptech.com/gb200-nvl72-nvidia-ai-rack-full-guide/) still apply.

[NVIDIA GB200 NVL72: A Full Guide to Liquid-Cooled AI Racks36 Grace + 72 Blackwell, NVLink scale-up, liquid-cooled rack facility planning, and honest limits for the generation before Vera Rubin.![](https://www.insidedeeptech.com/favicon.ico)Inside Deep Tech](https://www.insidedeeptech.com/gb200-nvl72-nvidia-ai-rack-full-guide/)

## Why NVL144 became NVL72, and why the numbers don't all match

If you followed the 2025 roadmap, you remember "Vera Rubin NVL144." NVIDIA's [Rubin CPX announcement](https://nvidianews.nvidia.com/news/nvidia-unveils-rubin-cpx-a-new-class-of-gpu-designed-for-massive-context-inference?ref=insidedeeptech.com) of September 9, 2025 described a Vera Rubin NVL144 CPX rack with **144** Rubin GPUs, **144** Rubin CPX GPUs and **36** Vera CPUs.

The 2026 materials describe the same CPU count with **72** Rubin GPUs ([NVIDIA](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com)), and each Rubin package contains two reticle-sized dies ([architecture blog](https://developer.nvidia.com/blog/inside-nvidia-rubin-gpu-architecture-powering-the-era-of-agentic-ai/?ref=insidedeeptech.com)). Inside Deep Tech reads this as a counting change, dies in 2025 and packages in 2026, rather than a smaller rack. NVIDIA hasn't published a formal explanation of the rename in the sources reviewed for this guide.

The spec mismatches are more practical. Here's where NVIDIA's own pages disagree:

| Metric                   | Product spec table                                                                              | Developer blogs and launch material                                                                                                                  |
| ------------------------ | ----------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- |
| Per-GPU HBM4 bandwidth   | [19.2 TB/s](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com)  | [Up to 22 TB/s](https://developer.nvidia.com/blog/inside-nvidia-rubin-gpu-architecture-powering-the-era-of-agentic-ai/?ref=insidedeeptech.com)       |
| Rack HBM4 bandwidth      | [1,400 TB/s](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com) | [1.6 PB/s](https://developer.nvidia.com/blog/how-the-nvidia-vera-rubin-platform-is-solving-agentic-ais-scale-up-problem/?ref=insidedeeptech.com)     |
| Per-GPU NVLink bandwidth | [3 TB/s](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com)     | [3.6 TB/s](https://developer.nvidia.com/blog/inside-the-nvidia-rubin-platform-six-new-chips-one-ai-supercomputer/?ref=insidedeeptech.com)            |
| Rack NVLink bandwidth    | [216 TB/s](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com)   | [260 TB/s](https://developer.nvidia.com/blog/nvidia-vera-rubin-pod-seven-chips-five-rack-scale-systems-one-ai-supercomputer/?ref=insidedeeptech.com) |

Source: [NVIDIA product page](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com), [Rubin GPU architecture blog](https://developer.nvidia.com/blog/inside-nvidia-rubin-gpu-architecture-powering-the-era-of-agentic-ai/?ref=insidedeeptech.com), [Vera Rubin platform blog](https://developer.nvidia.com/blog/inside-the-nvidia-rubin-platform-six-new-chips-one-ai-supercomputer/?ref=insidedeeptech.com), [Vera Rubin POD blog](https://developer.nvidia.com/blog/nvidia-vera-rubin-pod-seven-chips-five-rack-scale-systems-one-ai-supercomputer/?ref=insidedeeptech.com), [scale-up blog](https://developer.nvidia.com/blog/how-the-nvidia-vera-rubin-platform-is-solving-agentic-ais-scale-up-problem/?ref=insidedeeptech.com).

💡

Inside Deep Tech's take: assume the lower column until an OEM quotes otherwise in writing. The blog figures read like design targets and the spec table reads like the shipping SKU, and NVIDIA hasn't publicly explained the gap.

## HBM4: the memory that sets Rubin's ship rate

Rubin moves NVIDIA's data center GPUs to HBM4, which doubles the interface to **2,048** I/O terminals per stack ([SK hynix](https://news.skhynix.com/en/sk-hynix-completes-worlds-first-hbm4-development-and-readies-mass-production/?ref=insidedeeptech.com)). All three DRAM makers have announced HBM4 production:

| Supplier | Milestone                                                                                                                                                                                                                                                     | Published HBM4 spec                                                         |
| -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------- |
| SK hynix | [Development complete, mass production readied](https://news.skhynix.com/en/sk-hynix-completes-worlds-first-hbm4-development-and-readies-mass-production/?ref=insidedeeptech.com) (Sep 12, 2025)                                                              | 2,048 I/O, above 10 Gbps, over 40% better power efficiency                  |
| Samsung  | [Mass production and commercial shipments](https://news.samsungsemiconductor.com/global/samsung-ships-industry-first-commercial-hbm4-with-ultimate-performance-for-ai-computing/?ref=insidedeeptech.com) (Feb 12, 2026)                                       | 11.7 Gbps (up to 13 Gbps), up to 3.3 TB/s per stack, 24 to 36 GB at 12-high |
| Micron   | [High-volume production for Vera Rubin](https://investors.micron.com/news/press-release/2026/Micron-in-High-Volume-Production-of-HBM4-Designed-for-NVIDIA-Vera-Rubin-PCIe-Gen6-SSD-and-SOCAMM2-03-16-2026/default.aspx?ref=insidedeeptech.com) (Mar 16, 2026) | 36GB 12H, over 11 Gb/s, above 2.8 TB/s, 2.3x its HBM3E bandwidth            |

Source: [SK hynix Newsroom](https://news.skhynix.com/en/sk-hynix-completes-worlds-first-hbm4-development-and-readies-mass-production/?ref=insidedeeptech.com), [Samsung Semiconductor Newsroom](https://news.samsungsemiconductor.com/global/samsung-ships-industry-first-commercial-hbm4-with-ultimate-performance-for-ai-computing/?ref=insidedeeptech.com), [Micron investor release](https://investors.micron.com/news/press-release/2026/Micron-in-High-Volume-Production-of-HBM4-Designed-for-NVIDIA-Vera-Rubin-PCIe-Gen6-SSD-and-SOCAMM2-03-16-2026/default.aspx?ref=insidedeeptech.com).

NVIDIA's **288 GB** per GPU lines up with eight 36 GB 12-high stacks. That's Inside Deep Tech's arithmetic, not an NVIDIA disclosure; the [architecture blog](https://developer.nvidia.com/blog/inside-nvidia-rubin-gpu-architecture-powering-the-era-of-agentic-ai/?ref=insidedeeptech.com) confirms only 12-Hi stacks and the total.

The bigger issue is price. On the [Q2 fiscal 2027 call](https://s201.q4cdn.com/141608511/files/content%5Ffiles/TRANSCRIPT%5F-NVIDIA-Corp-NVDA-US-Q2-2027-Earnings-Call-26-August-2026-5%5F00-PM-ET.pdf?ref=insidedeeptech.com), NVIDIA's CFO said the company is "experiencing extreme pricing conditions in memory," and inventory rose to **$32 billion** as NVIDIA prepared for the Vera Rubin launch.

For how HBM stacks, base dies and packaging slots gate GPU supply, see Inside Deep Tech's [HBM full guide](https://www.insidedeeptech.com/high-bandwidth-memory-hbm-full-guide/).

[Read the HBM full guide](https://www.insidedeeptech.com/high-bandwidth-memory-hbm-full-guide/)

## Networking: NVLink 6, ConnectX-9, and co-packaged optics

Scale-up stays on copper inside the rack. NVLink 6 delivers **3.6 TB/s** per GPU across an all-to-all **72**\-GPU topology, with in-network compute for collectives, per NVIDIA's [CES release](https://nvidianews.nvidia.com/news/rubin-platform-ai-supercomputer?ref=insidedeeptech.com).

Scale-out moves to ConnectX-9 SuperNICs at **1.6 Tb/s** per GPU, and the switch tier moves to Spectrum-6 at **102.4 Tb/s** per chip with 200G PAM4 SerDes ([NVIDIA platform blog](https://developer.nvidia.com/blog/inside-the-nvidia-rubin-platform-six-new-chips-one-ai-supercomputer/?ref=insidedeeptech.com)). NVIDIA says Spectrum-X Ethernet Photonics, its co-packaged optics switch line, is in production and delivers **5x** better power efficiency than networks using traditional transceivers ([May 2026 release](https://nvidianews.nvidia.com/news/vera-rubin-full-production-agentic-ai-factory?ref=insidedeeptech.com)).

BlueField-4 DPUs offload networking, storage and security at up to **800Gb/s** ([NVIDIA](https://nvidianews.nvidia.com/news/vera-rubin-full-production-agentic-ai-factory?ref=insidedeeptech.com)). For which fabric does what, see Inside Deep Tech's [NVLink, InfiniBand, and UALink guide](https://www.insidedeeptech.com/nvlink-infiniband-ualink-ai-gpu-interconnect-full-guide/).

## Groq 3 LPX and the Rubin CPX question

The biggest change since the 2025 roadmap is the decode accelerator. NVIDIA's [GTC 2026 release](https://nvidianews.nvidia.com/news/nvidia-vera-rubin-platform?ref=insidedeeptech.com) added the Groq 3 LPX rack: **256** LPU processors with **128GB** of on-chip SRAM and **640 TB/s** of scale-up bandwidth, deployed alongside NVL72 and slated for availability in the second half of 2026.

NVL72 handles prefill and decode attention while LPUs run the latency-sensitive FFN decode loop via NVIDIA Dynamo's Attention-FFN Disaggregation, per NVIDIA's [scale-up blog](https://developer.nvidia.com/blog/how-the-nvidia-vera-rubin-platform-is-solving-agentic-ais-scale-up-problem/?ref=insidedeeptech.com). NVIDIA claims up to **35x** higher inference throughput per megawatt for trillion-parameter models with the combination ([GTC 2026 release](https://nvidianews.nvidia.com/news/nvidia-vera-rubin-platform?ref=insidedeeptech.com)). See Inside Deep Tech's [Groq LPU full guide](https://www.insidedeeptech.com/groq-lpu-nvidia-groq-3-lpx-full-guide/) for the LPU side.

Rubin CPX is the open question. NVIDIA announced it in September 2025 with **30 petaflops** of NVFP4, **128GB** of GDDR7 and end-of-2026 availability ([NVIDIA](https://nvidianews.nvidia.com/news/nvidia-unveils-rubin-cpx-a-new-class-of-gpu-designed-for-massive-context-inference?ref=insidedeeptech.com)). It was absent from the GTC 2026 keynote and slides ([Tom's Hardware](https://www.tomshardware.com/pc-components/gpus/nvidia-removes-rubin-cpx-accelerators-from-its-roadmap-groq-3-lpus-take-center-stage-as-cpx-is-removed?ref=insidedeeptech.com)), and later analyst reports of a redesigned CPX haven't been confirmed by NVIDIA. Treat any CPX-based plan as speculative until NVIDIA publishes a new spec.

## Performance claims, and what they're measured on

NVIDIA's headline claims are big, and every one has a footnote. On the [Vera Rubin NVL72 page](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com), NVIDIA claims:

- One-tenth the cost per million tokens versus GB200 NVL72, measured on Kimi-K2-Thinking at 32K/8K input and output lengths ([NVIDIA](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com)).
- Up to [10x more tokens per megawatt](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com) than GB200 NVL72 on the same model and settings.
- One-fourth the GPUs to train a 10T-parameter MoE model on 100T tokens in one month, labeled "projected performance subject to change" ([NVIDIA](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com)).
- [30x higher throughput per megawatt and 35x lower token cost](https://s201.q4cdn.com/141608511/files/content%5Ffiles/TRANSCRIPT%5F-NVIDIA-Corp-NVDA-US-Q2-2027-Earnings-Call-26-August-2026-5%5F00-PM-ET.pdf?ref=insidedeeptech.com) versus Grace Blackwell Ultra, stated on the Q2 fiscal 2027 earnings call.

Those claims use different baselines and NVIDIA-selected models. They're useful for direction, not for a purchase order.

Run your own model, sequence lengths and batch sizes before you sign a capacity plan.

## Facility planning: power, cooling, and what NVIDIA hasn't published

NVIDIA's spec table doesn't list a per-rack power figure for Vera Rubin NVL72 in the sources reviewed here. What it does publish is a reference design: about [40K Rubin GPUs](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com) in a 100 MW AI factory using NVIDIA DSX with MaxLPS.

NVIDIA says DSX MaxLPS can provision [up to 40% more GPUs](https://developer.nvidia.com/blog/inside-nvidia-rubin-gpu-architecture-powering-the-era-of-agentic-ai/?ref=insidedeeptech.com) in the same power budget, and the GTC release says DSX Max-Q enables [30% more AI infrastructure](https://nvidianews.nvidia.com/news/nvidia-vera-rubin-platform?ref=insidedeeptech.com) in a fixed-power data center. Both are provisioning claims, not lower chip power.

Longer term, Kyber racks are where power architecture breaks from today's halls. Inside Deep Tech's [800 VDC power distribution guide](https://www.insidedeeptech.com/800-vdc-power-distribution-ai-data-centers-full-guide/) explains why NVIDIA's Kyber generation pushes facilities toward 800-volt DC distribution.

🧊

Inside Deep Tech's take: don't size a Vera Rubin hall from a GB200 spreadsheet. Get the OEM's rack power, coolant flow and inlet temperature in writing, because NVIDIA's public spec table leaves the first one blank.

## Roadmap: Rubin Ultra NVL576, Kyber NVL144, and Feynman

NVIDIA's [POD blog](https://developer.nvidia.com/blog/nvidia-vera-rubin-pod-seven-chips-five-rack-scale-systems-one-ai-supercomputer/?ref=insidedeeptech.com) lays out the next two steps. Vera Rubin Ultra adds a two-layer all-to-all NVLink topology that links eight MGX NVL racks, each with **72** Rubin Ultra GPUs, into a **576**\-GPU domain using copper and direct optical connections.

Kyber is the next MGX rack design. It doubles the NVLink domain per rack to **144** GPUs and will first ship with Vera Rubin Ultra as a standalone NVL144, giving Rubin Ultra three scale-up options: NVL72, NVL144 and NVL576\. Eight Kyber racks then form an NVL1152 domain for the Feynman generation ([NVIDIA](https://developer.nvidia.com/blog/nvidia-vera-rubin-pod-seven-chips-five-rack-scale-systems-one-ai-supercomputer/?ref=insidedeeptech.com)).

So "NVL144" is coming back, this time as a Kyber rack with Rubin Ultra. NVIDIA gives no Rubin Ultra ship date in the sources reviewed here, so treat dates in secondary coverage as unconfirmed.

## Who should buy Vera Rubin now, and who should wait

Buy or reserve Vera Rubin NVL72 capacity when several of these are true:

- You train or serve large MoE models where [288 GB of HBM4 per GPU](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com) removes KV-cache offload.
- Your hall already runs liquid cooling and can supply [45°C coolant](https://developer.nvidia.com/blog/nvidia-vera-rubin-pod-seven-chips-five-rack-scale-systems-one-ai-supercomputer/?ref=insidedeeptech.com) at NVL72 density.
- You already run GB200 or GB300 racks and can reuse MGX logistics and CUDA tooling.

Wait, or stay on Blackwell, when any of these dominate:

- Your models fit comfortably in Blackwell memory and you're compute bound, not memory bound.
- You can't get allocation; NVIDIA expects [supply to remain a bottleneck](https://s201.q4cdn.com/141608511/files/content%5Ffiles/TRANSCRIPT%5F-NVIDIA-Corp-NVDA-US-Q2-2027-Earnings-Call-26-August-2026-5%5F00-PM-ET.pdf?ref=insidedeeptech.com) through fiscal 2028.
- Policy requires a second GPU source, which puts AMD Instinct on the shortlist.

## Honest limits

The real limits as of October 2026:

- Supply: NVIDIA expects [supply to remain a bottleneck](https://s201.q4cdn.com/141608511/files/content%5Ffiles/TRANSCRIPT%5F-NVIDIA-Corp-NVDA-US-Q2-2027-Earnings-Call-26-August-2026-5%5F00-PM-ET.pdf?ref=insidedeeptech.com) at least through the end of fiscal 2028.
- Memory cost: NVIDIA describes [extreme pricing conditions in memory](https://s201.q4cdn.com/141608511/files/content%5Ffiles/TRANSCRIPT%5F-NVIDIA-Corp-NVDA-US-Q2-2027-Earnings-Call-26-August-2026-5%5F00-PM-ET.pdf?ref=insidedeeptech.com), and that flows into rack prices.
- Spec ambiguity: NVIDIA's own pages list [19.2 TB/s](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com) or [22 TB/s](https://developer.nvidia.com/blog/inside-nvidia-rubin-gpu-architecture-powering-the-era-of-agentic-ai/?ref=insidedeeptech.com) per GPU, and 216 or 260 TB/s of NVLink per rack.
- Power: no per-rack power figure on NVIDIA's public spec table.
- Rubin CPX: status unconfirmed after its [absence from GTC 2026](https://www.tomshardware.com/pc-components/gpus/nvidia-removes-rubin-cpx-accelerators-from-its-roadmap-groq-3-lpus-take-center-stage-as-cpx-is-removed?ref=insidedeeptech.com).
- Groq 3 LPX is new, with availability in the [second half of 2026](https://nvidianews.nvidia.com/news/nvidia-vera-rubin-platform?ref=insidedeeptech.com) and no production track record yet.
- Performance multipliers are vendor projections on vendor-chosen models and baselines.

---

## Frequently asked questions

#### What is NVIDIA Vera Rubin?

NVIDIA Vera Rubin is NVIDIA's AI platform after Blackwell, built from seven chips including the Rubin GPU and Vera CPU. Its flagship system, [Vera Rubin NVL72](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com), links 72 Rubin GPUs and 36 Vera CPUs over NVLink 6.

#### When does Vera Rubin ship?

Production shipments began in August 2026, per NVIDIA's [Q2 fiscal 2027 earnings call](https://s201.q4cdn.com/141608511/files/content%5Ffiles/TRANSCRIPT%5F-NVIDIA-Corp-NVDA-US-Q2-2027-Earnings-Call-26-August-2026-5%5F00-PM-ET.pdf?ref=insidedeeptech.com). The [earnings release](https://nvidianews.nvidia.com/news/nvidia-announces-financial-results-for-second-quarter-fiscal-2027?ref=insidedeeptech.com) said racks were running at partners including CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure and Nebius.

#### Is Vera Rubin NVL144 the same as Vera Rubin NVL72?

For the first-generation rack, effectively yes. NVIDIA's [2025 materials](https://nvidianews.nvidia.com/news/nvidia-unveils-rubin-cpx-a-new-class-of-gpu-designed-for-massive-context-inference?ref=insidedeeptech.com) counted 144 Rubin GPUs alongside 36 Vera CPUs, while 2026 materials count 72 dual-die packages. The Kyber-based NVL144 coming with Rubin Ultra is a different, denser rack ([NVIDIA](https://developer.nvidia.com/blog/nvidia-vera-rubin-pod-seven-chips-five-rack-scale-systems-one-ai-supercomputer/?ref=insidedeeptech.com)).

#### How does Vera Rubin compare with GB200 NVL72?

NVIDIA claims one-tenth the cost per million tokens and up to 10x more tokens per megawatt versus GB200 NVL72 on Kimi-K2-Thinking, per its [product page](https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/?ref=insidedeeptech.com). Per-GPU HBM bandwidth rises 2.8x over Blackwell, per the [architecture blog](https://developer.nvidia.com/blog/inside-nvidia-rubin-gpu-architecture-powering-the-era-of-agentic-ai/?ref=insidedeeptech.com).

#### What happened to Rubin CPX?

Its status is unconfirmed. NVIDIA announced Rubin CPX in [September 2025](https://nvidianews.nvidia.com/news/nvidia-unveils-rubin-cpx-a-new-class-of-gpu-designed-for-massive-context-inference?ref=insidedeeptech.com) for end-of-2026 availability, but it didn't appear in the GTC 2026 keynote, [Tom's Hardware reported](https://www.tomshardware.com/pc-components/gpus/nvidia-removes-rubin-cpx-accelerators-from-its-roadmap-groq-3-lpus-take-center-stage-as-cpx-is-removed?ref=insidedeeptech.com), and NVIDIA now emphasizes Groq 3 LPX for low-latency decode.

#### Does Vera Rubin need liquid cooling?

Yes. NVIDIA says its third-generation MGX racks can be 100% liquid-cooled and are designed for a 45°C warm-water inlet, per its [Vera Rubin POD blog](https://developer.nvidia.com/blog/nvidia-vera-rubin-pod-seven-chips-five-rack-scale-systems-one-ai-supercomputer/?ref=insidedeeptech.com).

## What to do next

If you're planning 2027 capacity, the work this quarter is concrete. Get OEM rack power and coolant specs in writing, benchmark your own models against the conservative spec column, and lock HBM-heavy configurations early while [supply remains constrained](https://s201.q4cdn.com/141608511/files/content%5Ffiles/TRANSCRIPT%5F-NVIDIA-Corp-NVDA-US-Q2-2027-Earnings-Call-26-August-2026-5%5F00-PM-ET.pdf?ref=insidedeeptech.com).

The next decision point is Rubin Ultra. The facility you build for NVL72 now decides whether Kyber is an upgrade or a rebuild.

[Explore more Inside Deep Tech AI hardware guides](https://www.insidedeeptech.com/)