Google closed much of the hyperscaler-ASIC story for readers yesterday. AWS still runs a parallel silicon stack that most GPU shopping lists ignore until the invoice arrives.

This guide covers AWS Trainium and AWS Inferentia: the training-first and inference-first Neuron chips Amazon designs for EC2, SageMaker, and Bedrock-class workloads. You will get Trainium2 / NeuronCore-v3 architecture, Trn2 instances and UltraServers, Inferentia2 / Inf2 specs, NeuronLink and EFA scale-out, the Neuron SDK path (PyTorch, JAX, NKI), when these ASICs beat GPU clusters and when they lose, and an honest limits section. Numbers below come from AWS Neuron architecture docs and AWS EC2 product pages as of September 2026, not invented datasheets.

Key Takeaways

  • Trainium2 packs 8 NeuronCore-v3s, 96 GiB HBM, and 1,299 FP8 TFLOPS per chip.
  • trn2.48xlarge ties 16 chips for 20.8 FP8 PFLOPS and 1.5 TB HBM; UltraServers scale to 64 chips.
  • Inferentia2 is the inference sibling: 2 NeuronCore-v2s, 32 GiB HBM, up to 12 chips on Inf2.
  • Neuron SDK targets PyTorch and JAX; CUDA gravity and kernel depth still decide many teams.
  • These ASICs win on AWS price and capacity; they lose on odd kernels, scarce regions, and CUDA lock-in.
📌
Scope note: This is a named-thing Full Guide to AWS Trainium and Inferentia hardware and access models, not a glossary of "what is an ASIC," not a HubSpot category page, and not a how-to for every Neuron compiler flag. Peak FLOPs, HBM, and interconnect figures are AWS Neuron architecture and EC2 Trn2/Inf2 product values dated through September 2026. Customer quotes on AWS marketing pages (Anthropic Project Rainier, Bedrock latency modes) are vendor-published testimonials, not independent audits. Comparison context for Google TPU comes from Inside Deep Tech's TPU guide and Google Cloud docs cited there.

What AWS Trainium and Inferentia actually are

AWS Trainium is Amazon's purpose-built training accelerator family. Inferentia is the inference-optimized sibling. Both sit under the AWS Neuron software stack and ship as EC2 instance families (Trn1 / Trn2 and Inf1 / Inf2), not as cards you bolt into a random PCIe chassis in your lab.

That packaging choice matters. You rent topology, networking, and memory pooling as Amazon defined them. You do not assemble an open NVLink domain from mixed vendor parts. Inside Deep Tech's take is that this is closer to Google Cloud TPU than to a general GPU SKU: hyperscaler silicon with a closed interconnect story and a cloud-first software path.

As of September 2026, the generations that matter for new jobs are Trainium2 (NeuronCore-v3 on Trn2 / UltraServer) and Inferentia2 (NeuronCore-v2 on Inf2). First-generation Trainium and Inferentia remain relevant for migration math, but new capacity planning centers on the second wave.

AWS Neuron's Trainium2 architecture page is the primary source. Each Trainium2 chip has eight NeuronCore-v3 cores. Beginning with Trainium2, Neuron adds Logical NeuronCore Configuration (LNC), which lets you combine physical cores into larger logical cores when the model map wants fewer, fatter devices.

Per-chip compute and memory, from the Neuron Trainium2 hardware page:

  • 1,299 FP8 TFLOPS dense; 667 BF16 / FP16 / TF32 TFLOPS; 181 FP32 TFLOPS; 2,563 sparse TFLOPS across FP8/FP16/BF16/TF32.
  • 96 GiB device HBM at 2.9 TB/s bandwidth.
  • 3.5 TB/s DMA bandwidth with inline memory compression and decompression.
  • NeuronLink-v3 chip-to-chip interconnect at 1.28 TB/s per chip for collectives and memory pooling.
  • 16 CC-Cores for collective communication within and across instances.

Versus first-generation Trainium, Neuron publishes rough improvement factors of about 6.7× FP8, 3.4× BF16/FP16/TF32, 3.7× FP32, 3× HBM capacity, 3.6× HBM bandwidth, and 3.3× inter-chip interconnect bandwidth. Those are datasheet deltas, not measured tokens per second on your model.

Memory pressure still decides day-to-day wins. Large HBM helps batch size and KV cache, but poorly tiled kernels stall on bandwidth the same way they do on GPUs and TPUs. Inside Deep Tech's HBM full guide and TSMC CoWoS guide remain the packaging and supply companions: every serious accelerator, Trainium included, is gated by stacked DRAM and assembly slots.

Trn2 instances and UltraServers

Hardware only matters when Amazon exposes it as an EC2 shape. Neuron's Trn2 architecture page and the AWS Trn2 product page agree on the core topology.

A trn2.48xlarge (and the UltraServer-capable trn2u.48xlarge) puts 16 Trainium2 chips on one instance, tied with NeuronLink-v3. That yields about 20.8 FP8 PFLOPS, 10.7 PFLOPS at BF16/FP16/TF32, 1,536 GiB of device memory, 46.4 TB/s aggregate device bandwidth, 192 vCPUs, 2 TiB host memory, and 3.2 Tbps EFAv3 networking (Neuron Trn2 architecture table; AWS Trn2 product page).

A Trn2 UltraServer joins four trn2u.48xlarge instances so 64 Trainium2 chips share NeuronLink across the node. Neuron lists 83.2 FP8 PFLOPS, 42.8 PFLOPS at BF16/FP16/TF32, 6,144 GiB device memory, and 185.6 TB/s aggregate device bandwidth. AWS's product page also quotes about 12.8 Tbps EFAv3 for UltraServers and positions them for lower per-token latency on large models plus faster collective communication for model-parallel training.

Smaller shapes exist too. The product table lists trn2.3xlarge with a single Trainium2 chip and 96 GB accelerator memory for lighter jobs and bring-up.

SpecTrainium2 chiptrn2.48xlargeTrn2 UltraServer
NeuronCores8× v3128× v3512× v3
FP8 peak1,299 TFLOPS20.8 PFLOPS83.2 PFLOPS
BF16/FP16/TF32 peak667 TFLOPS10.7 PFLOPS42.8 PFLOPS
Device memory96 GiB1,536 GiB6,144 GiB
Device bandwidth2.9 TB/s46.4 TB/s185.6 TB/s
Chip-to-chip fabricNeuronLink-v3NeuronLink-v3 (intra)NeuronLink-v3 (intra + inter-instance)
Scale-out networkn/a3.2 Tbps EFAv3Up to 12.8 Tbps EFAv3 (product page)

Source: AWS Neuron Trainium2 and Trn2 architecture docs; Amazon EC2 Trn2 instances product page (accessed September 2026). Sparse FLOPs (2,563 TFLOPS per chip; 41 / 164 PFLOPS at instance / UltraServer) are separate from dense FP8 and must not be conflated in capacity planning.

Inferentia2 and Inf2: the inference sibling

If Trainium2 is the training and dual-use hammer, Inferentia2 is the inference-optimized chisel. Neuron's Inferentia2 page says each chip has two NeuronCore-v2 cores delivering 380 INT8 TOPS, 190 FP16/BF16/cFP8/TF32 TFLOPS, and 47.5 FP32 TFLOPS, plus 32 GiB HBM at 820 GiB/s, with NeuronLink-v2 for multi-chip sharding.

Inf2 instance shapes (Neuron Inf2 architecture table):

InstanceInferentia2 chipsDevice memory (GiB)FP8/FP16/BF16/TF32 TFLOPSNetwork (Gbps)
inf2.xlarge132190Up to 15
inf2.8xlarge132190Up to 25
inf2.24xlarge61921,14050
inf2.48xlarge123842,280100

Source: AWS Neuron Inf2 / Inferentia2 architecture docs and the Amazon EC2 Inf2 product page. AWS marketing claims Inf2 delivers up to about 4× throughput and up to about 10× lower latency versus Inf1 on selected workloads. Treat those as vendor claims tied to specific models and batch shapes.

Practical split for operators: use Inf2 when the job is steady-state serving, cost-sensitive tokens, and models that compile cleanly under Neuron. Use Trn2 when you need training, heavy fine-tuning, or UltraServer-scale memory pooling for frontier-size inference. Mixing both in one fleet is normal. Pretending they are interchangeable FLOPs puddles is not.

Trainium2's scale-up story is NeuronLink-v3 inside the instance or UltraServer. Scale-out across UltraClusters rides Elastic Fabric Adapter (EFAv3) on the Nitro System. That is a different vocabulary from NVLink domains plus InfiniBand or UALink open scale-up.

Inside Deep Tech's NVLink / InfiniBand / UALink guide still applies as the conceptual map: scale-up is the high-bandwidth chip-to-chip domain; scale-out is the rack-to-rack fabric. On AWS Neuron silicon, NeuronLink is the scale-up column and EFA is the scale-out column. On NVIDIA stacks, NVLink (and increasingly NVLink-C2C) and InfiniBand / Ethernet RoCE play those roles.

PCIe remains the host attach path for many accelerators. It is not the GPU-to-GPU or Trainium-to-Trainium happy path. See the PCIe 6.0 full guide for why host links and accelerator fabrics must stay in separate BOM columns.

Software path: Neuron SDK, PyTorch, JAX, and NKI

Hardware without a compiler path is scrap. AWS ships the Neuron SDK with native hooks for PyTorch and JAX, Deep Learning AMIs and containers, and integrations that AWS lists for SageMaker, EKS, ECS, ParallelCluster, Batch, Hugging Face Optimum Neuron, Lightning, Ray, and others.

For performance engineers, Neuron Kernel Interface (NKI) exposes a Python / Triton-like path to the ISA so you can write custom kernels instead of waiting for every fused op to land in the stock compiler. That is real power. It is also real work. CUDA's library depth (cuDNN, Transformer Engine, FlashAttention variants, vendor-tuned MoE kernels) remains the gravity well most teams already paid for.

AWS product copy claims "train and deploy without changing a single line of code" for some PyTorch paths. Inside Deep Tech's reading of operator reality is narrower: many Hugging Face and Lightning models port with modest changes, but production latency SLOs, custom attention, and exotic MoE routing still need Neuron-aware engineering. Budget for that engineering, or do not put Trainium on the critical path.

🎯
Inside Deep Tech's take: Trainium2 and Inferentia2 are serious hyperscaler ASICs with published architecture numbers, not press-release vapor. They win when your stack already lives on AWS, your models compile cleanly under Neuron, and price or GPU scarcity hurts more than CUDA convenience. They lose when your differentiation is custom CUDA kernels, multi-cloud portability, or a toolchain that assumes NVIDIA libraries on day one. Peak PFLOPS on a datasheet never paid an SRE's pager.

When Trainium / Inferentia beat GPUs, and when they do not

AWS markets Trn2 as offering about 30–40% better price performance than EC2 P5e / P5en GPU instances, and Trn2 as about 3× more energy efficient than Trn1. Those are vendor claims. Use them as a negotiation starting point, then measure your own tokens per dollar and joules per token.

Where Neuron silicon tends to win:

  • AWS-native training and serving with Neuron-supported model families (Llama-class, diffusion, many Hugging Face hubs models AWS cites).
  • Cost-sensitive inference on Inf2 when the graph compiles and batching stays healthy.
  • UltraServer-scale memory pooling when a single GPU node would force awkward pipeline parallel cuts.
  • Capacity hedges when GPU UltraClusters are sold out in the regions you need.

Where Neuron silicon tends to lose:

  • Workloads that depend on bleeding-edge CUDA libraries or research kernels not yet mapped to NKI.
  • Teams that must stay portable across GCP TPUs, Azure Maia, on-prem DGX, and AWS without a second software track.
  • Odd control flow, dynamic shapes that thrash the compiler, or tiny batches that leave systolic-style engines idle.
  • Regions or instance types where Trn2 / Inf2 simply are not available at the volume you need.

Anthropic's Project Rainier and Bedrock latency-optimized Claude modes appear on AWS's Trn2 marketing page as customer testimonials tied to large Trainium2 clusters. Treat them as evidence that frontier labs will run on Neuron silicon at scale, not as a guarantee that your 7B fine-tune will match Claude's engineering budget.

Honest limits

Name the downside or the guide is marketing with footnotes.

  • Published FLOPs are peak. Sparse modes are not dense FP8. Do not paste UltraServer sparse PFLOPS into a dense capacity model.
  • NeuronLink topologies are Amazon's, not your open fabric. You cannot casually mix Trainium with UALink or NVLink domains.
  • Software maturity trails CUDA for exotic kernels even when PyTorch "just works" for common graphs.
  • UltraServer networking numbers differ slightly between Neuron tables and the marketing product page; pin your planning doc to one primary source and re-check before a purchase order.
  • Liquid cooling, power delivery, and hall design still gate dense AI racks whether the accelerator says Trainium or GPU. See Inside Deep Tech's liquid cooling and 800 VDC guides when the facility is the real bottleneck.
Google TPU: A Full Guide to Tensor Processing Units
Companion hyperscaler ASIC guide: Ironwood/TPU7x, Trillium/v6e, Pod/Slice/ICI, JAX/XLA, honest limits versus GPUs and Trainium2.

FAQ

What is the difference between AWS Trainium and Inferentia?

Trainium is AWS's training-first accelerator family; Inferentia is the inference-optimized sibling. Both use the Neuron SDK. As of September 2026, new capacity centers on Trainium2 (Trn2 / UltraServer) and Inferentia2 (Inf2).

How many FLOPS does Trainium2 deliver?

Per AWS Neuron docs, one Trainium2 chip peaks at 1,299 FP8 TFLOPS and 667 BF16/FP16/TF32 TFLOPS. A trn2.48xlarge with 16 chips is rated at 20.8 FP8 PFLOPS; a 64-chip UltraServer at 83.2 FP8 PFLOPS.

What is a Trn2 UltraServer?

It is four trn2u.48xlarge instances linked so 64 Trainium2 chips share NeuronLink, pooling about 6 TiB of device memory and 83.2 FP8 PFLOPS for large training and low-latency frontier inference.

Should I use Inf2 or Trn2 for LLM inference?

Start with Inf2 for cost-sensitive serving that compiles cleanly. Move to Trn2 or UltraServers when model size, KV cache, or latency needs exceed Inf2's 12-chip / 384 GB ceiling or when you already train on Trainium.

Does Neuron support PyTorch and JAX?

Yes. AWS documents native PyTorch and JAX integration in the Neuron SDK, plus Hugging Face Optimum Neuron and other framework hooks. Custom kernels go through NKI when stock ops are not enough.

How does Trainium2 compare to Google TPU?

Both are hyperscaler ASICs with closed interconnects and cloud-first software. TPUs center on MXU systolic arrays and ICI Pods; Trainium2 centers on NeuronCore-v3, NeuronLink, and EFA. Pick the cloud you already operate, then measure.

When should I stay on NVIDIA GPUs instead?

Stay on GPUs when CUDA libraries, multi-cloud portability, or research kernels dominate your roadmap, or when Trn2/Inf2 capacity is missing in your regions. ASICs win on price and AWS-native scale, not on universal software coverage.


Monday's move is simple. If your fleet already lives on AWS and GPU quotes keep slipping, stand up a Neuron bring-up path on a single Trainium2 or Inferentia2 shape, compile one production graph end to end, and measure tokens per dollar against your current P5 or G-class baseline. Datasheet PFLOPS can wait until that experiment finishes.