Google TPU and AWS Trainium closed two sides of the hyperscaler-ASIC triangle for Inside Deep Tech readers this week. Azure still runs a third path that most GPU shopping lists skip until Copilot latency or Foundry token economics show up on the invoice.

This guide covers Microsoft Maia: Azure's custom AI accelerator family, from Maia 100 (Hot Chips 2024) to Maia 200 (announced January 2026, technical paper August 2026). You will get architecture numbers from Microsoft's Hot Chips slides, Azure Blog, and the Maia 200 arXiv paper; how Ethernet scale-up fits the stack; software paths (PyTorch, Triton, Maia API / NPL); where Cobalt Arm CPUs sit as host context; when Maia beats GPU clusters and when it does not; and an honest limits section. Numbers below are vendor-published as of September 2026, not invented datasheets.

Key Takeaways

  • Maia 100: TSMC N5, ~820 mm², 64 GB HBM2E at 1.8 TB/s, 700W design / 500W provision TDP.
  • Maia 200: TSMC 3nm, 216 GB HBM3e at 7 TB/s, ~10.1 PFLOPS FP4 / 5.1 PFLOPS FP8, 750W TDP.
  • Both generations use Ethernet-based scale-up/scale-out rather than NVLink-style proprietary fabrics.
  • Software centers on PyTorch, Triton, and Maia-specific APIs; CUDA gravity still decides many teams.
  • Maia wins inside Azure first-party inference; it loses when you need rentable multi-cloud GPU portability.
📌
Scope note: This is a named-thing Full Guide to Microsoft Maia hardware and access models, not a glossary of "what is an ASIC," not a HubSpot category page, and not a how-to for every Maia compiler flag. Peak FLOPs, HBM, and interconnect figures come from Microsoft Hot Chips 2024 (Maia 100), the Azure Blog (April 2024), Scott Guthrie's Microsoft Blog (January 26, 2026), and arXiv:2608.24664 (August 25, 2026). Comparison claims versus Trainium3 or Google TPU7x on Microsoft marketing pages are vendor claims, not independent audits. Cobalt appears only as Azure Arm host-CPU context. Companion context for Google TPU and AWS Trainium comes from Inside Deep Tech's prior Full Guides.

What Microsoft Maia actually is

Microsoft Maia is Azure's first-party AI accelerator line. It is not a GPU you buy from NVIDIA or AMD, and it is not the same product as Azure's rented ND MI300X or ND H100 GPU VM families. Maia is silicon Microsoft designs for its own cloud, co-optimized with racks, liquid cooling, Ethernet fabrics, and the software that serves Copilot, Foundry, and OpenAI models hosted on Azure.

That packaging choice matters. Like Google Cloud TPU and AWS Trainium / Inferentia, Maia is hyperscaler silicon with a closed topology story and a cloud-first software path. Inside Deep Tech's take is that Maia sits closer to those ASICs than to a general GPU SKU you assemble into an open NVLink domain.

As of September 2026, two generations matter for capacity planning: Maia 100 (announced around Ignite 2023, detailed at Hot Chips 2024, production OpenAI / Copilot-class workloads) and Maia 200 (inference-focused follow-on, Microsoft Blog January 26, 2026; arXiv paper August 25, 2026). Maia 200 is already in Microsoft's US Central fleet near Des Moines, with US West 3 near Phoenix next.

Maia 100 architecture: N5, HBM2E, and Ethernet

The primary hardware source for Maia 100 is Microsoft's Hot Chips 2024 talk (Sherry Xu and Chandru Ramakrishnan). The Azure Blog from April 3, 2024 fills in systems context.

Per Hot Chips 2024, each Maia 100 chip is roughly 820 mm² on TSMC N5 in CoWoS-S packaging, with 64 GB HBM2E at 1.8 TB/s bandwidth and about 500 MB of L1/L2 SRAM. Design TDP is 700 W; production provision TDP is 500 W for the inference envelopes Microsoft described.

Peak dense tensor throughput on the Hot Chips slide is quoted in POPS: 3 POPS at 6-bit, 1.5 POPS at 9-bit, and 0.8 POPS at BF16. Host attach is PCIe Gen5 x8 (32 GB/s). Backend network bandwidth is 600 GB/s via 12×400GbE. The Azure Blog separately cites a fully custom Ethernet-based protocol with 4.8 Tbps aggregate bandwidth per accelerator for scale.

Inside the SoC, Hot Chips shows 16 clusters of 4 tiles each, with tensor units (TTU), vector engines (TVP), tile DMA, and tile control processors on a high-bandwidth mesh NoC. Microsoft positioned Maia 100 as the first production implementation of the Microscaling (MX) data format (Open Compute Project v1.0, co-published with AMD, Arm, Intel, Meta, NVIDIA, and Qualcomm).

Memory pressure still decides day-to-day wins. HBM2E at 64 GB was deliberately behind the HBM3e race so Microsoft would not fight NVIDIA and AMD for the hottest stacks. Inside Deep Tech's HBM full guide and TSMC CoWoS guide remain the packaging companions: every serious accelerator is gated by stacked DRAM and assembly slots.

SpecMaia 100 (Hot Chips 2024)
Process / dieTSMC N5, ~820 mm²
PackageTSMC CoWoS-S
HBM64 GB HBM2E @ 1.8 TB/s
On-chip SRAM (L1/L2)~500 MB
Peak dense tensor6-bit 3 POPS; 9-bit 1.5 POPS; BF16 0.8 POPS
Host linkPCIe Gen5 x8 (32 GB/s)
Backend network600 GB/s (12×400GbE); Azure Blog also cites 4.8 Tbps aggregate
TDP700 W design / 500 W provision
Topology note16 clusters × 4 tiles; Ethernet RoCE-like; AES-GCM

Source: Microsoft Hot Chips 2024 "Inside Maia 100" slides; Azure Blog "Azure Maia for the era of AI" (April 3, 2024). POPS are peak dense tensor figures as published on the Hot Chips slide, not measured tokens per second.

Maia 200: inference-first SDLA at 3 nm

Maia 200 is the second generation, framed explicitly for efficient large-scale inference. Scott Guthrie's Microsoft Blog (January 26, 2026) and the Microsoft technical paper on arXiv (2608.24664, August 25, 2026) are the primary sources.

Published chip-level numbers:

  • TSMC 3 nm process; more than 140 billion transistors; near-reticle monolithic die (~26×33 mm) in CoWoS-S; ~75×75 mm package.
  • 216 GB HBM3e at 7 TB/s; 272 MB on-chip SRAM (Microsoft Blog).
  • arXiv peak: 10,145 TFLOPS FP4 and 5,072 TFLOPS FP8 within a 750 W TDP (about 13.3 / 6.7 TFLOPS/W).
  • Microsoft Blog marketing summary: over 10 PFLOPS FP4 and over 5 PFLOPS FP8 inside a 750 W SoC TDP.
  • 2.8 TB/s bidirectional dedicated scale-up bandwidth; clusters of up to 6,144 accelerators on all-Ethernet networking.
  • Four Maia accelerators per tray fully connected with direct non-switched links; same protocols for intra-rack and inter-rack.

Architecturally, Maia 200 implements what Microsoft calls a Software Defined Locally Accessed Dataflow Architecture (SDLA): explicit programming of data-movement engines and specialized memories, rather than classical GPU thread-centric SIMT as the primary contract. That is a systems claim. Measure your models against it; do not treat the acronym as a free lunch.

Deployment status as of the January 2026 announcement: Maia 200 was live in US Central near Des Moines, with US West 3 near Phoenix next. Microsoft said it would serve GPT-5.2-class OpenAI models, Microsoft Foundry, Microsoft 365 Copilot, and Microsoft Superintelligence workloads (synthetic data and reinforcement learning). Those are first-party and partner deployments, not a public "ND Maia" SKU table.

DimensionMaia 100Maia 200Trainium2 (chip)Google TPU context
ProcessTSMC N5TSMC 3nmAWS Neuron docsCloud TPU gen-dependent
HBM64 GB HBM2E @ 1.8 TB/s216 GB HBM3e @ 7 TB/s96 GiB @ 2.9 TB/sSee TPU guide
Peak (published)BF16 0.8 POPS; MX 6/9-bit peaks10,145 TFLOPS FP4; 5,072 TFLOPS FP81,299 FP8 TFLOPSSee TPU guide / Cloud docs
TDP700W design / 500W provision750W SoCNot one-line public chip TDPPod/Slice power varies
Scale-up fabricEthernet (custom RoCE-like)All-Ethernet two-tier; 2.8 TB/s bi-dirNeuronLink-v3ICI
Primary accessAzure first-party / servicesFirst-party + SDK previewEC2 Trn2 / UltraServerCloud TPU VMs / Pods
SoftwarePyTorch, Triton, Maia APIPyTorch, Triton, NPL, simulatorNeuron SDK / NKIJAX/XLA, PyTorch

Sources: Hot Chips 2024 (Maia 100); Microsoft Blog and arXiv:2608.24664 (Maia 200); AWS Neuron Trainium2 docs and Inside Deep Tech's Trainium / TPU guides for comparison columns. Do not treat cross-vendor peak FLOPs as interchangeable capacity units.

Maia's interconnect story is Ethernet, not NVLink. Hot Chips described a custom RoCE-like protocol with AES-GCM, unified scale-up and scale-out, 4800 Gbps all-gather / scatter-reduce, and 1200 Gbps any-to-any on Maia 100. Maia 200 doubles down with a two-tier all-Ethernet design and Microsoft Collective Communication Library (MCCL) style software on top.

That choice aligns with Microsoft's role in the Ultra Ethernet Consortium. It is also a different vocabulary from NVLink domains plus InfiniBand, or from open UALink scale-up. Inside Deep Tech's NVLink / InfiniBand / UALink guide remains the conceptual map: scale-up is the high-bandwidth chip-to-chip domain; scale-out is the rack-to-rack fabric. On Maia, Ethernet carries both roles with AI-specific transport and endpoint hardware.

Optical I/O and copper still sit under the same power and reach constraints every cluster faces. When copper scale-up runs out of reach, see Inside Deep Tech's Optical I/O full guide. Maia does not erase physics; it chooses Ethernet as the control plane and data plane language.

Cobalt: Arm host CPU context, not the hero

Azure Cobalt is Microsoft's custom Arm CPU line (Cobalt 100 based on Arm Neoverse N2 at 3.4 GHz; Cobalt 200 discussed in later Azure infrastructure posts). Microsoft Learn documents Cobalt 100 VM series such as Dplsv6 / Dpsv6 / Epsv6 for general-purpose and memory-optimized Linux workloads.

Cobalt is host and general compute context in Azure's silicon story. It is not the AI accelerator. Do not conflate Cobalt vCPU catalogs with Maia FLOPs. Operators evaluating AI silicon care about Maia memory, collectives, and compiler paths; Cobalt matters when you size the Arm hosts sitting beside those accelerators.

Software path: PyTorch, Triton, Maia API, and the SDK preview

Hardware without a compiler path is scrap. For Maia 100, Hot Chips and the Azure Blog describe a Maia SDK with PyTorch integration, ONNX Runtime hooks, Triton for portable kernels, and a lower-level Maia API for maximum control. The stack includes a cuBLAS-like kernel library, NCCL-like collectives, debugger / profiler / visualizer tools, and a maia-smi style device utility.

For Maia 200, Microsoft opened an SDK preview at announcement: Triton compiler, PyTorch support, low-level NPL programming, plus a simulator and cost calculator. Sign-up is for developers, startups, and academics optimizing models ahead of broader hardware access. That is not the same as spinning up a public Azure VM SKU today.

CUDA's library depth remains the gravity well most teams already paid for. Triton helps portability across accelerators when kernels stay in the portable subset. Exotic attention, MoE routing, and research ops still need Maia-aware engineering. Budget for that engineering, or do not put Maia on the critical path of a customer-facing latency SLO.

🎯
Inside Deep Tech's take: Maia is serious hyperscaler silicon with published architecture numbers across two generations, not press-release vapor. It wins when your inference economics live inside Azure first-party services, your graphs compile cleanly under Triton / Maia tooling, and GPU scarcity or token cost hurts more than CUDA convenience. It loses when your differentiation is custom CUDA kernels, multi-cloud portability, or a need to rent accelerators tomorrow as a self-managed VM shape. Peak PFLOPS on a blog post never paid an SRE's pager.

When Maia beats GPUs (and Trainium / TPU), and when it does not

Microsoft markets Maia 200 as about 30% better performance per dollar than the latest generation hardware in its fleet, with FP4 performance claims versus AWS Trainium3 and FP8 claims versus Google's seventh-generation TPU. Those are vendor claims on Microsoft's blog. Use them as a negotiation and research starting point, then measure your own tokens per dollar and joules per token.

Where Maia tends to win:

  • Azure-native inference for Copilot, Foundry, and OpenAI models Microsoft already co-designs onto Maia.
  • Ethernet-centric clusters where proprietary NVLink domains are unavailable or unwanted.
  • Cost and capacity hedges when GPU SKUs are sold out but Microsoft can schedule Maia capacity for first-party or partner work.
  • Teams willing to invest in Triton / Maia SDK bring-up for a durable Azure inference track.

Where Maia tends to lose:

  • Workloads that need a public, self-serve accelerator VM tomorrow (Trn2 / Inf2 and Cloud TPU are clearer rent paths as of September 2026).
  • Teams locked to bleeding-edge CUDA libraries or research kernels not mapped to Triton / NPL.
  • Multi-cloud portability mandates that forbid a second Azure-specific software track.
  • Odd control flow, dynamic shapes that thrash compilers, or tiny batches that leave dataflow engines idle.

Who this is for (and not for)

This guide is for infra engineers, investors, and operators who evaluate Azure AI silicon alongside NVIDIA GPUs, Google TPUs, and AWS Neuron chips. If you already run production on Azure GPU SKUs and want a factual Maia baseline, you are the reader.

This guide is not for buyers hunting a Maia SKU in the Azure portal price list this week. If your constraint is "rent 8 accelerators by Friday," start with ND-class GPU VMs, Cloud TPU, or Trn2 / Inf2, and treat Maia as a strategic Azure first-party path until Microsoft publishes general customer VM access.

Honest limits

Name the downside or the guide is marketing with footnotes.

  • Published FLOPs are peak. FP4 and FP8 peaks are not interchangeable with BF16 training capacity or measured tokens per second.
  • Customer access remains first-party / partner / SDK-preview heavy. Do not assume a public ND Maia catalog exists because Maia 200 is in production for Copilot.
  • Ethernet fabrics are Microsoft's, not your open NVLink or UALink domain. Mixing Maia with arbitrary GPU scale-up is not a weekend project.
  • Software maturity trails CUDA for exotic kernels even when PyTorch "just works" for common graphs.
  • Regional capacity (US Central first for Maia 200) and liquid-cooling / power delivery still gate dense AI racks whether the accelerator says Maia or GPU.
AWS Trainium and Inferentia: A Full Guide to Neuron ASICs
Companion hyperscaler ASIC guide: Trainium2 / Inferentia2, NeuronLink/EFA, Neuron SDK, honest limits versus GPUs and TPUs.

FAQ

What is Microsoft Maia?

Maia is Microsoft's custom Azure AI accelerator family. Maia 100 (Hot Chips 2024) and Maia 200 (January 2026) are designed for Azure-scale AI workloads, especially inference for Copilot, Foundry, and OpenAI models hosted on Azure.

What are the Maia 100 specs?

Per Hot Chips 2024: ~820 mm² on TSMC N5 with CoWoS-S, 64 GB HBM2E at 1.8 TB/s, ~500 MB L1/L2, 12×400GbE backend networking, 700 W design TDP and 500 W provision TDP, with peak dense tensor figures of 3 / 1.5 / 0.8 POPS at 6-bit / 9-bit / BF16.

How does Maia 200 differ from Maia 100?

Maia 200 is inference-focused on TSMC 3 nm with 216 GB HBM3e at 7 TB/s, about 10.1 PFLOPS FP4 and 5.1 PFLOPS FP8 at 750 W (arXiv:2608.24664), all-Ethernet scale-up to thousands of accelerators, and an SDLA programming model. Maia 100 was the first-generation N5 part with HBM2E.

Can customers rent Maia like an Azure GPU VM?

As of September 2026, Maia capacity is described mainly as Microsoft first-party and partner deployment plus an SDK preview. That is different from publicly documented EC2 Trn2/Inf2 or Google Cloud TPU rental models. Check Azure for any newer SKU announcements before planning a purchase order.

What software runs on Maia?

Microsoft documents PyTorch integration, Triton for kernels, ONNX Runtime hooks on Maia 100, and for Maia 200 an SDK preview with Triton, PyTorch, NPL, a simulator, and a cost calculator. CUDA libraries do not run natively.

How does Maia compare to Google TPU and AWS Trainium?

All three are hyperscaler ASICs with closed interconnects and cloud-first software. TPUs center on MXU/ICI Pods; Trainium on NeuronCore/NeuronLink/EFA; Maia on Ethernet fabrics and Azure first-party inference. Pick the cloud you already operate, then measure.

What is Azure Cobalt relative to Maia?

Cobalt is Microsoft's custom Arm CPU for general Azure VMs (for example Cobalt 100 Neoverse N2 hosts). It is host compute context, not the AI accelerator. Maia is the AI ASIC; Cobalt is the Arm CPU line.

When should I stay on NVIDIA GPUs instead?

Stay on GPUs when CUDA libraries, multi-cloud portability, or immediate self-serve VM rental dominate your roadmap, or when Maia access is unavailable for your tenant. ASICs win on Azure-native inference economics, not on universal software coverage.


Monday's move is simple. If your inference already lives on Azure and GPU quotes keep slipping, inventory whether your models are candidates for Maia's Triton / PyTorch path, request SDK preview access if you are eligible, and measure one production graph's tokens per dollar against your current ND-class GPU baseline. Datasheet PFLOPS can wait until that experiment finishes.