NVIDIA still owns most AI rack mindshare. Buyers who need another OAM option keep landing on the same question: can AMD Instinct MI350 actually fill a training or inference row when CUDA is not the only stack on the RFP?

The short answer is that MI350 is a real CDNA 4 generation, not a rebadge. AMD's MI350X product page lists 288 GB HBM3E, 8 TB/s peak memory bandwidth, and a 1000W typical board power on a passive OAM, with launch dated 6/12/2025. The 8-GPU MI350X Platform puts 2.3 TB of HBM3E on a UBB 2.0 baseboard.

This full guide walks through what AMD Instinct MI350 actually is (MI350X, MI355X, MI350P), the CDNA 4 chiplet math, how Infinity Fabric scale-up differs from NVLink rack domains, where ROCm stands, OEM packaging from Supermicro, and the honest limits versus NVIDIA Blackwell racks. Pair it with Inside Deep Tech's GB200 NVL72 guide when you need the NVIDIA liquid-cooled rack counterpart.

Key Takeaways

⚠️
AMD Instinct MI350 is an OAM generation first, a rack narrative second. Spec-sheet FLOPS matter, but the buying decision still turns on ROCm day-zero readiness, HBM allocation, and whether an 8-GPU Infinity Fabric island is enough for the model.

What AMD Instinct MI350 actually is

Per AMD's MI350 Series overview and the ROCm MI350 microarchitecture docs, the series spans three form factors:

  • MI350X: 1000W air-cooled passive OAM with eight XCDs, two IODs, and 288 GB HBM3E.
  • MI355X: same memory capacity and CDNA 4 die plan at 1400W for direct liquid-cooled denser clocks.
  • MI350P: full-height PCIe card with four XCDs, 144 GB HBM3E at 4 TB/s, 600W (configurable to 450W).

CDNA 4 is a chiplet machine. ROCm docs describe eight accelerator complex dies (XCDs) on TSMC N3P plus two I/O dies on TSMC N6, tied with on-package Infinity Fabric and eight stacks of 12-Hi HBM3E (36 GB per stack). That is deliberate process splitting: logic density where it pays, I/O and memory controllers on the mature node.

Active compute lands at 256 compute units (16,384 stream processors, 1,024 matrix cores) with a 2200 MHz peak engine clock and 185 billion transistors on the MI350X product sheet.

Specs that matter for buyers

Sparse versus dense FLOPS footnotes matter. Treat marketing multipliers as workload-specific, and prefer AMD's published matrix, vector, and sparsity columns over reseller summaries.

MetricMI350X (AMD / ROCm)Why it matters
ArchitectureCDNA4 (TSMC 3nm | 6nm)Process split for XCDs vs IODs
HBM288 GB HBM3E at 8 TB/sKV-cache and MoE weight headroom
TBP1000W passive OAMAir-row vs DLC planning
MXFP4 / MXFP6 matrix9.2 PFLOPsLow-precision inference path
OCP-FP8 matrix4.6 PFLOPs (sparse 9.2)Main FP8 training/inference band
FP16 matrix2.3 PFLOPs (sparse 4.6)Dense half-precision baseline
FP64 / FP3272.1 / 144.2 TFLOPsHPC and mixed scientific loads
Infinity Cache256 MBMemory-side cache on IODs

Source: AMD Instinct MI350X product page; ROCm MI350 Series microarchitecture.

Memory packaging still gates volume. HBM3E stacks and advanced packaging capacity decide how many OAM modules OEMs can ship; see Inside Deep Tech's HBM full guide and TSMC CoWoS packaging guide.

The 8-GPU MI350X Platform on UBB 2.0

AMD's MI350X Platform page describes an industry-standard UBB 2.0 compatible board with 8 MI350X OAMs, dimensions 417mm x 553mm, and 2.3 TB total HBM3E. Per-OAM bandwidth stays at 8.0 TB/s. Aggregate bi-directional peer-to-peer I/O is listed at 1,194.8 GB/s.

Platform-level matrix peaks on that page include 73.8 PFLOPs MXFP4, 36.9 PFLOPs OCP-FP8 dense (73.8 PFLOPS sparse), and 18.5 PFLOPs FP16 matrix (36.9 PFLOPs sparse). HPC columns list 1.2 PFLOPs FP32 and 576.8 TFLOPs FP64 for the full eight-GPU board.

ROCm's node description matches the topology: each GPU keeps one PCIe Gen 5 x16 host link and seven Infinity Fabric links at 38.4 Gbps, with more than 1 TB/s of aggregate communication bandwidth per GPU inside the fully connected eight-GPU island (ROCm MI350 docs).

💡
Inside Deep Tech's take: treat Infinity Fabric on MI350 as an 8-GPU scale-up island, not a substitute for NVIDIA's 72-GPU NVLink domain. Confusing those products is how RFPs under-buy scale-out networking and over-promise single-node miracles.
NVIDIA GB200 NVL72: A Full Guide to Liquid-Cooled AI Racks
36 Grace + 72 Blackwell, 130 TB/s NVLink, 120-135 kW liquid-cooled racks, and honest facility limits.

Generation jump: MI300X to MI325X to MI350X

ROCm's MI300 / MI350 workload optimization guide is the cleanest official side-by-side. Capacity and bandwidth climb hard; CU count actually drops as matrix throughput and LDS grow.

FeatureMI300XMI325XMI350X
ArchitectureCDNA3CDNA3CDNA4
Memory192 GB HBM3256 GB HBM3E288 GB HBM3E
Bandwidth5.3 TB/s6 TB/s8 TB/s
Active CUs304304256
Max power750W1000W1000W
FP8 (dense)2.6 PF FNUZ2.61 PF4.6 PF OCP
MXFP4 / MXFP6N/AN/A9.2 PF
LDS per CU64 KB64 KB160 KB

Source: ROCm MI300/MI350 architecture comparison; AMD MI325X product page; AMD MI350X product page.

Two software gotchas sit in those rows. FP8 moves from the FNUZ variant on MI300-class parts to OCP FP8 on MI350, so quantized checkpoints are not drop-in. TF32 matrix hardware disappears in favor of software emulation via BF16 on CDNA 4, while BF16 matrix throughput rises (ROCm notes). FP64 matrix rate is also lower on MI350X than on MI300X, which HPC buyers should re-benchmark rather than assume a free upgrade.

MXFP4, MXFP6, and why low precision is the product bet

CDNA 4's new hardware story is micro-scaling. AMD lists peak 9.2 PFLOPs for both MXFP4 and MXFP6 matrix, alongside 4.6 PFLOPs MXFP8 and OCP-FP8. ROCm documents OCP MX formats with a shared exponent across blocks of 32 elements (workload optimization guide).

That is the same industry direction NVIDIA markets with NVFP4 on Blackwell. The buyer question is never who printed the bigger PFLOPS cell. It is whether your serving stack, quantization pipeline, and accuracy gates actually land on MX or FP8 paths without weeks of kernel work.

AMD's product-page comparison versus B200 SXM5 180GB leans on memory (288 GB vs 180 GB), bandwidth (8.0 vs 7.7 TB/s), sparse OCP-FP8 (9.2 vs 9), and MXFP6 versus FP6 Tensor (9.2 vs 4.5). Those are peak theoretical columns with AMD footnotes. Re-run your own MoE and dense checkpoints before treating any of them as a purchase order.

MI350 keeps the fully connected 8-GPU node topology from the prior Instinct generation. ROCm states Infinity Fabric links run at 38.4 Gbps (up from 32 Gbps on MI300), with P2P ring aggregate bandwidth at 1,075.2 GB/s and total peak aggregate I/O at 1,203.2 GB/s. AMD's platform page publishes 1,194.8 GB/s bi-directional peer-to-peer I/O for the UBB board.

That is strong for an 8-GPU OAM island. It is not a 72-GPU copper NVLink domain. Inside Deep Tech's NVLink, InfiniBand, and UALink guide and UALink full guide cover why scale-up and scale-out are different BOM columns. AMD participates in the open UALink narrative for multi-vendor scale-up; production MI350 nodes today still ship on Infinity Fabric inside the UBB.

For NVIDIA's rack-scale answer, read the GB200 NVL72 full guide: 72 Blackwell GPUs, 130 TB/s NVLink domain bandwidth, and roughly 120–135 kW liquid-cooled racks. Different product, different facility ask.

OEM reality: Supermicro H14 air and liquid SKUs

On June 12, 2025, Supermicro announced H14 systems with MI350 Series GPUs in both air-cooled and liquid-cooled form (Supermicro IR release). The H14 8-GPU datasheet lists the 8U air-cooled MI350X system AS-8126GS-TNMR and the 4U liquid-cooled MI355X system AS-4126GS-NMR-LCC.

OEM packaging details that matter for RFPs: OAM modules on UBB 2.0, 2.3 TB HBM3E per node, dual 5th Gen AMD EPYC hosts, up to 9 TB DDR5-6000, and 400-Gbps networking dedicated to each GPU. Supermicro frames the platform as a seamless upgrade path from MI325X with ROCm day-zero continuity.

Liquid-cooled MI355X rows inherit the usual CDU, leak-detection, and facility-water questions. Use Inside Deep Tech's data center liquid cooling full guide when the SKU sheet says DLC but the hall still thinks in CRAH CFM.

ROCm, PyTorch, and vLLM: the software gate

Hardware without a serving path is scrap silicon. AMD's ROCm stack is the official path for Instinct, with PyTorch, hipBLASLt, Triton, RCCL, and vLLM called out heavily in the MI350 workload optimization docs. The docs walk TunableOp GEMM search, torch.compile / Inductor on AMD GPUs, RCCL eight-GPU collective guidance, and vLLM V1 tuning for MI350X / MI355X.

That documentation density is progress. It is not CUDA parity. Teams should budget porting time for custom CUDA kernels, NCCL assumptions baked into training scripts, and ISV certifications that still list NVIDIA first. Day-zero claims on a press release are not the same as your MoE checkpoint converging on ROCm without a war room.

Partitioning is part of the software story too. ROCm documents compute partitions of 1, 2, 4, or 8 XCDs and memory modes NPS1 (full 288 GB interleaved) versus NPS2 (two 144 GB pools) on MI350X / MI355X (microarchitecture docs). Inference multi-tenancy and MIG-like isolation patterns need those knobs validated before you promise density to finance.

When MI350 wins (and when it does not)

Choose AMD Instinct MI350 when several of these are true:

  • You need 288 GB HBM3E per GPU for large KV caches or multi-tenant inference without jumping to a 72-GPU NVLink rack.
  • Your software path is ROCm-ready (PyTorch, vLLM, Triton) or you can fund the port.
  • You want OCP UBB / OAM second-source leverage versus a single GPU vendor.
  • HPC mixed with AI needs strong FP64/FP32 columns; AMD quotes 72.1 TFLOPs FP64 on MI350X.
  • You can take MI350X air-cooled UBB density or fund MI355X DLC for 1400W parts.

Stay on NVIDIA (or wait) when any of these dominate:

  • The model truly needs a 72-GPU NVLink domain and CUDA/NCCL lock-in is already paid for (see GB200 NVL72).
  • Critical kernels or ISVs remain CUDA-only with no ROCm roadmap.
  • Procurement cannot secure HBM3E / packaging allocation for Instinct volume.
  • Ops cannot run an 8-GPU OAM plus 400GbE scale-out design cleanly.

Supply chain: HBM3E, packaging, and lead times

Every CDNA 4 OAM still depends on HBM stacks and advanced packaging slots. Inside Deep Tech covers those bottlenecks in the HBM guide and CoWoS guide. Quote MI350 as a platform (OAM + UBB + host + NICs + cooling), not a unit GPU fantasy price.

MI325X remains the bridge SKU for buyers who need Instinct capacity before MI350 volumes clear. AMD's MI325X page lists 256 GB HBM3E at 6 TB/s and 1000W peak TBP on CDNA 3. Supermicro explicitly sells MI350 as an upgrade from that platform.

What a real MI350 deployment checklist looks like

Field teams should treat an 8-GPU UBB node like a mini cluster, not a single GPU card:

  • Confirm row power and cooling for 1000W MI350X air OAMs or 1400W MI355X DLC parts before PO.
  • Validate UBB 2.0 / OAM seating, Infinity Fabric link training, and PCIe Gen 5 host attach.
  • Bring up 400-Gbps per-GPU NIC paths and RCCL all-reduce before model burn-in.
  • Pin ROCm, PyTorch, and vLLM versions; run TunableOp / Inductor passes documented in ROCm optimization guides.
  • Re-quantize or re-export FP8/MX checkpoints for OCP formats; do not assume FNUZ MI300 artifacts load cleanly.
  • Soak under sustained inference and collective loads; watch thermal margins on air 8U versus liquid 4U SKUs.

Density math: eight fat OAMs still change the row

Eight 1000W OAMs are already an 8 kW GPU drawer before hosts, NICs, and storage. MI355X at 1400W each pushes the GPU tray alone toward 11.2 kW. That is not GB200 NVL72's 120 kW+ rack, but it is far past casual air-cooled hobby density. Plan bus bars, breaker coordination, and hot-aisle containment accordingly, and escalate to DLC when the SKU sheet says MI355X.

Honest limits

Inside Deep Tech will not pretend MI350 erases NVIDIA's software gravity or invents a 72-GPU AMD NVLink clone. The real limits are:

  • Scale-up world size stays eight GPUs on Infinity Fabric inside the UBB.
  • ROCm maturity and ISV coverage still trail CUDA for many production kernels.
  • HBM3E and packaging allocation can gate ship dates as hard as silicon.
  • FP8 format change (FNUZ to OCP) and TF32 software emulation create migration friction from MI300.
  • FP64 matrix throughput is lower on MI350X than MI300X; HPC must re-bench.
  • Multi-rack scale-out still needs Ethernet or InfiniBand design, not wishful Infinity Fabric extension.

FAQ

What is AMD Instinct MI350?

AMD Instinct MI350 is AMD's CDNA 4 accelerator series. The flagship MI350X is a 1000W OAM with 288 GB HBM3E and 8 TB/s bandwidth, launched 6/12/2025. MI355X is the 1400W liquid-cooled twin; MI350P is a PCIe card with 144 GB HBM3E.

How does MI350X compare to MI325X?

Per ROCm architecture tables, MI350X moves to CDNA 4 with 288 GB HBM3E at 8 TB/s versus MI325X's 256 GB at 6 TB/s on CDNA 3. Matrix low-precision throughput rises and LDS grows to 160 KB, while active CU count drops from 304 to 256.

How does MI350 compare to NVIDIA Blackwell / GB200?

AMD publishes peak comparisons versus B200 SXM5 on the MI350X page (memory, bandwidth, FP8/MX peaks). Rack-scale NVLink domains such as GB200 NVL72 are a different product class; see Inside Deep Tech's GB200 NVL72 guide.

What interconnect does MI350 use for GPU-to-GPU?

Inside the node, seven Infinity Fabric links per GPU at 38.4 Gbps form a fully connected 8-GPU island. Scale-out uses host networking (OEM sheets emphasize 400 Gbps per GPU). Open UALink is the longer-term multi-vendor scale-up narrative.

Is liquid cooling required for MI350?

MI350X ships as a 1000W passive OAM for air-cooled UBB systems (for example Supermicro 8U). MI355X is specified at 1400W for direct liquid cooling. Cooling is SKU-dependent, not optional on the DLC part.

What software stack runs on MI350?

AMD ROCm is the primary stack, with PyTorch, vLLM, Triton, hipBLASLt, and RCCL covered in ROCm optimization docs. Expect porting effort wherever custom CUDA kernels dominate.

How much HBM does an 8-GPU MI350X platform have?

AMD's platform page lists 2.3 TB total HBM3E across eight MI350X OAMs, with 8.0 TB/s bandwidth per OAM.

When should a buyer pick MI350 over waiting for more NVIDIA allocation?

When ROCm readiness, 288 GB HBM3E per GPU, OAM second-source leverage, and an 8-GPU island match the workload, and when HBM/packaging supply for Instinct is actually available. If you need a 72-GPU NVLink domain, evaluate GB200 NVL72 on its own facility terms.