NVIDIA still owns most AI rack mindshare. Buyers who need another OAM option keep landing on the same question: can AMD Instinct MI350 actually fill a training or inference row when CUDA is not the only stack on the RFP?
The short answer is that MI350 is a real CDNA 4 generation, not a rebadge. AMD's MI350X product page lists 288 GB HBM3E, 8 TB/s peak memory bandwidth, and a 1000W typical board power on a passive OAM, with launch dated 6/12/2025. The 8-GPU MI350X Platform puts 2.3 TB of HBM3E on a UBB 2.0 baseboard.
This full guide walks through what AMD Instinct MI350 actually is (MI350X, MI355X, MI350P), the CDNA 4 chiplet math, how Infinity Fabric scale-up differs from NVLink rack domains, where ROCm stands, OEM packaging from Supermicro, and the honest limits versus NVIDIA Blackwell racks. Pair it with Inside Deep Tech's GB200 NVL72 guide when you need the NVIDIA liquid-cooled rack counterpart.
Key Takeaways
- AMD Instinct MI350X is CDNA 4 OAM silicon: 288 GB HBM3E at 8 TB/s, 1000W TBP, launch 6/12/2025.
- Headline matrix peaks include 9.2 PFLOPs MXFP4/MXFP6 and 4.6 PFLOPs OCP-FP8 (9.2 PFLOPs with structured sparsity).
- An 8-GPU MI350X Platform aggregates 2.3 TB HBM3E with 1,194.8 GB/s peer-to-peer I/O on UBB 2.0.
- MI355X is the 1400W direct liquid-cooled OAM twin; MI350P is a 144 GB / 4 TB/s PCIe card for mainstream servers.
- Honest limits: 8-GPU Infinity Fabric scale-up (not 72-GPU NVLink), ROCm vs CUDA maturity, CoWoS/HBM allocation, and OCP-FP8 / MX datatype migration from MI300.
What AMD Instinct MI350 actually is
Per AMD's MI350 Series overview and the ROCm MI350 microarchitecture docs, the series spans three form factors:
- MI350X: 1000W air-cooled passive OAM with eight XCDs, two IODs, and 288 GB HBM3E.
- MI355X: same memory capacity and CDNA 4 die plan at 1400W for direct liquid-cooled denser clocks.
- MI350P: full-height PCIe card with four XCDs, 144 GB HBM3E at 4 TB/s, 600W (configurable to 450W).
CDNA 4 is a chiplet machine. ROCm docs describe eight accelerator complex dies (XCDs) on TSMC N3P plus two I/O dies on TSMC N6, tied with on-package Infinity Fabric and eight stacks of 12-Hi HBM3E (36 GB per stack). That is deliberate process splitting: logic density where it pays, I/O and memory controllers on the mature node.
Active compute lands at 256 compute units (16,384 stream processors, 1,024 matrix cores) with a 2200 MHz peak engine clock and 185 billion transistors on the MI350X product sheet.
Specs that matter for buyers
Sparse versus dense FLOPS footnotes matter. Treat marketing multipliers as workload-specific, and prefer AMD's published matrix, vector, and sparsity columns over reseller summaries.
| Metric | MI350X (AMD / ROCm) | Why it matters |
|---|---|---|
| Architecture | CDNA4 (TSMC 3nm | 6nm) | Process split for XCDs vs IODs |
| HBM | 288 GB HBM3E at 8 TB/s | KV-cache and MoE weight headroom |
| TBP | 1000W passive OAM | Air-row vs DLC planning |
| MXFP4 / MXFP6 matrix | 9.2 PFLOPs | Low-precision inference path |
| OCP-FP8 matrix | 4.6 PFLOPs (sparse 9.2) | Main FP8 training/inference band |
| FP16 matrix | 2.3 PFLOPs (sparse 4.6) | Dense half-precision baseline |
| FP64 / FP32 | 72.1 / 144.2 TFLOPs | HPC and mixed scientific loads |
| Infinity Cache | 256 MB | Memory-side cache on IODs |
Source: AMD Instinct MI350X product page; ROCm MI350 Series microarchitecture.
Memory packaging still gates volume. HBM3E stacks and advanced packaging capacity decide how many OAM modules OEMs can ship; see Inside Deep Tech's HBM full guide and TSMC CoWoS packaging guide.
The 8-GPU MI350X Platform on UBB 2.0
AMD's MI350X Platform page describes an industry-standard UBB 2.0 compatible board with 8 MI350X OAMs, dimensions 417mm x 553mm, and 2.3 TB total HBM3E. Per-OAM bandwidth stays at 8.0 TB/s. Aggregate bi-directional peer-to-peer I/O is listed at 1,194.8 GB/s.
Platform-level matrix peaks on that page include 73.8 PFLOPs MXFP4, 36.9 PFLOPs OCP-FP8 dense (73.8 PFLOPS sparse), and 18.5 PFLOPs FP16 matrix (36.9 PFLOPs sparse). HPC columns list 1.2 PFLOPs FP32 and 576.8 TFLOPs FP64 for the full eight-GPU board.
ROCm's node description matches the topology: each GPU keeps one PCIe Gen 5 x16 host link and seven Infinity Fabric links at 38.4 Gbps, with more than 1 TB/s of aggregate communication bandwidth per GPU inside the fully connected eight-GPU island (ROCm MI350 docs).
Generation jump: MI300X to MI325X to MI350X
ROCm's MI300 / MI350 workload optimization guide is the cleanest official side-by-side. Capacity and bandwidth climb hard; CU count actually drops as matrix throughput and LDS grow.
| Feature | MI300X | MI325X | MI350X |
|---|---|---|---|
| Architecture | CDNA3 | CDNA3 | CDNA4 |
| Memory | 192 GB HBM3 | 256 GB HBM3E | 288 GB HBM3E |
| Bandwidth | 5.3 TB/s | 6 TB/s | 8 TB/s |
| Active CUs | 304 | 304 | 256 |
| Max power | 750W | 1000W | 1000W |
| FP8 (dense) | 2.6 PF FNUZ | 2.61 PF | 4.6 PF OCP |
| MXFP4 / MXFP6 | N/A | N/A | 9.2 PF |
| LDS per CU | 64 KB | 64 KB | 160 KB |
Source: ROCm MI300/MI350 architecture comparison; AMD MI325X product page; AMD MI350X product page.
Two software gotchas sit in those rows. FP8 moves from the FNUZ variant on MI300-class parts to OCP FP8 on MI350, so quantized checkpoints are not drop-in. TF32 matrix hardware disappears in favor of software emulation via BF16 on CDNA 4, while BF16 matrix throughput rises (ROCm notes). FP64 matrix rate is also lower on MI350X than on MI300X, which HPC buyers should re-benchmark rather than assume a free upgrade.
MXFP4, MXFP6, and why low precision is the product bet
CDNA 4's new hardware story is micro-scaling. AMD lists peak 9.2 PFLOPs for both MXFP4 and MXFP6 matrix, alongside 4.6 PFLOPs MXFP8 and OCP-FP8. ROCm documents OCP MX formats with a shared exponent across blocks of 32 elements (workload optimization guide).
That is the same industry direction NVIDIA markets with NVFP4 on Blackwell. The buyer question is never who printed the bigger PFLOPS cell. It is whether your serving stack, quantization pipeline, and accuracy gates actually land on MX or FP8 paths without weeks of kernel work.
AMD's product-page comparison versus B200 SXM5 180GB leans on memory (288 GB vs 180 GB), bandwidth (8.0 vs 7.7 TB/s), sparse OCP-FP8 (9.2 vs 9), and MXFP6 versus FP6 Tensor (9.2 vs 4.5). Those are peak theoretical columns with AMD footnotes. Re-run your own MoE and dense checkpoints before treating any of them as a purchase order.
Infinity Fabric scale-up versus NVLink and UALink
MI350 keeps the fully connected 8-GPU node topology from the prior Instinct generation. ROCm states Infinity Fabric links run at 38.4 Gbps (up from 32 Gbps on MI300), with P2P ring aggregate bandwidth at 1,075.2 GB/s and total peak aggregate I/O at 1,203.2 GB/s. AMD's platform page publishes 1,194.8 GB/s bi-directional peer-to-peer I/O for the UBB board.
That is strong for an 8-GPU OAM island. It is not a 72-GPU copper NVLink domain. Inside Deep Tech's NVLink, InfiniBand, and UALink guide and UALink full guide cover why scale-up and scale-out are different BOM columns. AMD participates in the open UALink narrative for multi-vendor scale-up; production MI350 nodes today still ship on Infinity Fabric inside the UBB.
For NVIDIA's rack-scale answer, read the GB200 NVL72 full guide: 72 Blackwell GPUs, 130 TB/s NVLink domain bandwidth, and roughly 120–135 kW liquid-cooled racks. Different product, different facility ask.
OEM reality: Supermicro H14 air and liquid SKUs
On June 12, 2025, Supermicro announced H14 systems with MI350 Series GPUs in both air-cooled and liquid-cooled form (Supermicro IR release). The H14 8-GPU datasheet lists the 8U air-cooled MI350X system AS-8126GS-TNMR and the 4U liquid-cooled MI355X system AS-4126GS-NMR-LCC.
OEM packaging details that matter for RFPs: OAM modules on UBB 2.0, 2.3 TB HBM3E per node, dual 5th Gen AMD EPYC hosts, up to 9 TB DDR5-6000, and 400-Gbps networking dedicated to each GPU. Supermicro frames the platform as a seamless upgrade path from MI325X with ROCm day-zero continuity.
Liquid-cooled MI355X rows inherit the usual CDU, leak-detection, and facility-water questions. Use Inside Deep Tech's data center liquid cooling full guide when the SKU sheet says DLC but the hall still thinks in CRAH CFM.
ROCm, PyTorch, and vLLM: the software gate
Hardware without a serving path is scrap silicon. AMD's ROCm stack is the official path for Instinct, with PyTorch, hipBLASLt, Triton, RCCL, and vLLM called out heavily in the MI350 workload optimization docs. The docs walk TunableOp GEMM search, torch.compile / Inductor on AMD GPUs, RCCL eight-GPU collective guidance, and vLLM V1 tuning for MI350X / MI355X.
That documentation density is progress. It is not CUDA parity. Teams should budget porting time for custom CUDA kernels, NCCL assumptions baked into training scripts, and ISV certifications that still list NVIDIA first. Day-zero claims on a press release are not the same as your MoE checkpoint converging on ROCm without a war room.
Partitioning is part of the software story too. ROCm documents compute partitions of 1, 2, 4, or 8 XCDs and memory modes NPS1 (full 288 GB interleaved) versus NPS2 (two 144 GB pools) on MI350X / MI355X (microarchitecture docs). Inference multi-tenancy and MIG-like isolation patterns need those knobs validated before you promise density to finance.
When MI350 wins (and when it does not)
Choose AMD Instinct MI350 when several of these are true:
- You need 288 GB HBM3E per GPU for large KV caches or multi-tenant inference without jumping to a 72-GPU NVLink rack.
- Your software path is ROCm-ready (PyTorch, vLLM, Triton) or you can fund the port.
- You want OCP UBB / OAM second-source leverage versus a single GPU vendor.
- HPC mixed with AI needs strong FP64/FP32 columns; AMD quotes 72.1 TFLOPs FP64 on MI350X.
- You can take MI350X air-cooled UBB density or fund MI355X DLC for 1400W parts.
Stay on NVIDIA (or wait) when any of these dominate:
- The model truly needs a 72-GPU NVLink domain and CUDA/NCCL lock-in is already paid for (see GB200 NVL72).
- Critical kernels or ISVs remain CUDA-only with no ROCm roadmap.
- Procurement cannot secure HBM3E / packaging allocation for Instinct volume.
- Ops cannot run an 8-GPU OAM plus 400GbE scale-out design cleanly.
Supply chain: HBM3E, packaging, and lead times
Every CDNA 4 OAM still depends on HBM stacks and advanced packaging slots. Inside Deep Tech covers those bottlenecks in the HBM guide and CoWoS guide. Quote MI350 as a platform (OAM + UBB + host + NICs + cooling), not a unit GPU fantasy price.
MI325X remains the bridge SKU for buyers who need Instinct capacity before MI350 volumes clear. AMD's MI325X page lists 256 GB HBM3E at 6 TB/s and 1000W peak TBP on CDNA 3. Supermicro explicitly sells MI350 as an upgrade from that platform.
What a real MI350 deployment checklist looks like
Field teams should treat an 8-GPU UBB node like a mini cluster, not a single GPU card:
- Confirm row power and cooling for 1000W MI350X air OAMs or 1400W MI355X DLC parts before PO.
- Validate UBB 2.0 / OAM seating, Infinity Fabric link training, and PCIe Gen 5 host attach.
- Bring up 400-Gbps per-GPU NIC paths and RCCL all-reduce before model burn-in.
- Pin ROCm, PyTorch, and vLLM versions; run TunableOp / Inductor passes documented in ROCm optimization guides.
- Re-quantize or re-export FP8/MX checkpoints for OCP formats; do not assume FNUZ MI300 artifacts load cleanly.
- Soak under sustained inference and collective loads; watch thermal margins on air 8U versus liquid 4U SKUs.
Density math: eight fat OAMs still change the row
Eight 1000W OAMs are already an 8 kW GPU drawer before hosts, NICs, and storage. MI355X at 1400W each pushes the GPU tray alone toward 11.2 kW. That is not GB200 NVL72's 120 kW+ rack, but it is far past casual air-cooled hobby density. Plan bus bars, breaker coordination, and hot-aisle containment accordingly, and escalate to DLC when the SKU sheet says MI355X.
Honest limits
Inside Deep Tech will not pretend MI350 erases NVIDIA's software gravity or invents a 72-GPU AMD NVLink clone. The real limits are:
- Scale-up world size stays eight GPUs on Infinity Fabric inside the UBB.
- ROCm maturity and ISV coverage still trail CUDA for many production kernels.
- HBM3E and packaging allocation can gate ship dates as hard as silicon.
- FP8 format change (FNUZ to OCP) and TF32 software emulation create migration friction from MI300.
- FP64 matrix throughput is lower on MI350X than MI300X; HPC must re-bench.
- Multi-rack scale-out still needs Ethernet or InfiniBand design, not wishful Infinity Fabric extension.
FAQ
What is AMD Instinct MI350?
AMD Instinct MI350 is AMD's CDNA 4 accelerator series. The flagship MI350X is a 1000W OAM with 288 GB HBM3E and 8 TB/s bandwidth, launched 6/12/2025. MI355X is the 1400W liquid-cooled twin; MI350P is a PCIe card with 144 GB HBM3E.
How does MI350X compare to MI325X?
Per ROCm architecture tables, MI350X moves to CDNA 4 with 288 GB HBM3E at 8 TB/s versus MI325X's 256 GB at 6 TB/s on CDNA 3. Matrix low-precision throughput rises and LDS grows to 160 KB, while active CU count drops from 304 to 256.
How does MI350 compare to NVIDIA Blackwell / GB200?
AMD publishes peak comparisons versus B200 SXM5 on the MI350X page (memory, bandwidth, FP8/MX peaks). Rack-scale NVLink domains such as GB200 NVL72 are a different product class; see Inside Deep Tech's GB200 NVL72 guide.
What interconnect does MI350 use for GPU-to-GPU?
Inside the node, seven Infinity Fabric links per GPU at 38.4 Gbps form a fully connected 8-GPU island. Scale-out uses host networking (OEM sheets emphasize 400 Gbps per GPU). Open UALink is the longer-term multi-vendor scale-up narrative.
Is liquid cooling required for MI350?
What software stack runs on MI350?
AMD ROCm is the primary stack, with PyTorch, vLLM, Triton, hipBLASLt, and RCCL covered in ROCm optimization docs. Expect porting effort wherever custom CUDA kernels dominate.
How much HBM does an 8-GPU MI350X platform have?
AMD's platform page lists 2.3 TB total HBM3E across eight MI350X OAMs, with 8.0 TB/s bandwidth per OAM.
When should a buyer pick MI350 over waiting for more NVIDIA allocation?
When ROCm readiness, 288 GB HBM3E per GPU, OAM second-source leverage, and an 8-GPU island match the workload, and when HBM/packaging supply for Instinct is actually available. If you need a 72-GPU NVLink domain, evaluate GB200 NVL72 on its own facility terms.



