Google TPU, AWS Trainium, Microsoft Maia, and Cerebras Wafer-Scale Engine cover training-heavy and hyperscaler ASIC bets. Groq answers a different question: what if the chip is built only for deterministic, SRAM-first token generation, and what happens when that architecture becomes NVIDIA's decode accelerator inside Vera Rubin?

This guide covers the Groq Language Processing Unit (LPU) and NVIDIA Groq 3 LPX. You will get the four LPU design principles from Groq's own explainers; rack, tray, and chip numbers from NVIDIA's March 2026 architecture blog; how Attention-FFN Disaggregation and Dynamo split work with Vera Rubin NVL72; production and Artificial Analysis context from August 2026; how LPUs compare with GPU clusters and with Cerebras WSE, Google TPU, Microsoft Maia, and AWS Trainium; and an honest limits section. Numbers below are vendor-published as of September 2026 context, not invented datasheets.

Key Takeaways

  • LPU = SRAM-first, compiler-scheduled inference silicon; not a general training GPU.
  • NVIDIA Groq 3 LPX: 256 LPUs, 128 GB SRAM, 40 PB/s, 315 PFLOPS FP8 (rack).
  • AFD pairs Rubin GPUs (prefill/attention) with LPX (FFN/MoE decode).
  • Aug 2026: LPX in full production; ~3,400 tok/s on Gemma 4 31B @ 100K ctx (Artificial Analysis via NVIDIA).
  • Wins on interactive decode latency; loses when HBM capacity, training, or CUDA flexibility dominate.
📌
Scope note: This is a named-thing Full Guide to Groq LPU architecture and NVIDIA Groq 3 LPX as of September 2026, not a glossary of "what is inference," not a HubSpot category page, and not a how-to for every compiler flag. Rack FLOPs, SRAM, and bandwidth figures come from NVIDIA's Inside Groq 3 LPX technical blog (March 16, 2026). Token-rate and "35x / 10x" claims are NVIDIA or Artificial Analysis figures cited by NVIDIA, not independent audits by Inside Deep Tech. Companion context for other non-GPU AI silicon comes from Inside Deep Tech's Cerebras, TPU, Trainium, and Maia Full Guides.

What an LPU is (and why it is not a GPU)

A Language Processing Unit is Groq's name for an inference-focused processor built around four design principles the company published on March 7, 2025: software-first control, a programmable assembly-line (streaming) architecture, deterministic compute and networking, and on-chip memory.

GPUs were born for massively parallel graphics and later bent toward AI. They keep large model weights in off-chip HBM, schedule work dynamically at runtime, and win when flexibility and ecosystem depth matter. Groq's bet is narrower. Inference for large language models is mostly linear algebra with a predictable dataflow once a model is compiled. If software can schedule every memory move and every compute step down to the clock cycle, you can strip the caches, arbiters, and jitter that GPUs carry for generality.

That is the LPU pitch. Put SRAM next to compute. Let the compiler own the schedule. Stream tensors through functional units like an assembly line, within a chip and across chips, without waiting on a hardware scheduler to resolve contention.

Groq's own comparison language puts on-chip SRAM bandwidth "upwards of 80 TB/s" against GPU off-chip HBM around "eight TB/s," and claims architectural energy efficiency up to about 10x versus GPUs for inference. Treat those as vendor framing from the LPU explained post, then measure your workload. Inside Deep Tech's HBM full guide remains the right companion when you compare stacked DRAM next to a GPU with SRAM tiled inside an LPU die. They are not interchangeable BOM columns.

Four LPU design principles that still matter in 2026

Groq's March 2025 explainer is still the cleanest primary description of the architecture that later appears inside NVIDIA's Vera Rubin platform.

Software-first and a model-independent compiler

Groq says it designed the compiler architecture before it touched the first chip. The goal is to put the developer (and the compiler) in control of utilization instead of writing model-specific GPU kernels for every new network. The LPU path accepts workloads from common frameworks, maps them across one or many chips, and emits a program that already contains data-movement schedules.

Programmable assembly line

Function units sit on software-controlled "conveyor belts." Each step receives instructions that say where to read, what to compute, and where to write. The same streaming model extends chip-to-chip, so a multi-LPU system looks like a longer assembly line rather than a hub-and-spoke of caches and routers.

Deterministic compute and networking

Every step is intended to be predictable to the clock cycle. Contention for bandwidth and compute is designed out so the schedule does not thrash at runtime. That determinism is also what later NVIDIA power features (Preemptive Power and Clock Period Synthesis) exploit on Groq 3 LPX.

On-chip SRAM as the working set

Weights, activations, and other hot state live in on-chip SRAM for the active partition. That removes the round trip to HBM for the latency-critical path, at the cost of tiny per-chip capacity compared with a modern GPU HBM stack. Large models are therefore a multi-chip, rack-scale problem, not a single-package HBM problem.

NVIDIA Groq 3 LPX: the Vera Rubin decode accelerator

On March 16, 2026, NVIDIA published Inside NVIDIA Groq 3 LPX, positioning LPX as a rack-scale, low-latency inference accelerator co-designed with Vera Rubin NVL72. Rubin GPUs remain the general-purpose workhorse for training and for high-throughput prefill and attention. LPX is the specialized engine for fast, predictable token generation in interactive and agentic regimes.

NVIDIA's footer language on the August 24, 2026 production announcement is important for buyers: "Groq and LPU are used under license from Groq, Inc." Treat LPX as NVIDIA platform silicon that carries Groq's LPU architecture under license, not as a synonym for GroqCloud. Groq's inference cloud remains a separate product path and, per the same NVIDIA PR, plans to be among early LPX adopters.

Rack, tray, and chip specs (vendor numbers)

At rack scale, NVIDIA documents 256 interconnected Groq 3 LPU accelerators (LP30 chips) with these headline figures:

SpecificationNVIDIA Groq 3 LPX (rack)
AI inference compute (FP8)315 PFLOPS
Total SRAM capacity128 GB
On-chip SRAM bandwidth40 PB/s
Scale-up density256 chips
Scale-up bandwidth640 TB/s
Compute trays32 liquid-cooled 1U trays (8 LPUs each)

Source: NVIDIA Technical Blog, "Inside NVIDIA Groq 3 LPX" (March 16, 2026). Peak FLOPs and bandwidth are vendor architecture numbers, not third-party lab measurements.

Each 1U tray integrates eight LPU modules plus host and fabric expansion logic in a cableless, liquid-cooled design. Per tray, NVIDIA lists 4 GB on-chip SRAM, 1.2 PB/s SRAM bandwidth, up to 256 GB DRAM via fabric expansion logic, up to 128 GB DRAM via the host CPU, 9.6 PFLOPS FP8, and 20 TB/s scale-up bandwidth.

At the chip level, the NVIDIA Groq 3 LPU (described as the seventh chip of the Vera Rubin platform) puts 500 MB of compiler-managed SRAM in a flat MEM block, uses 320-byte vectors as the unit of work, and exposes specialized MXM (matrix), VXM (vector / activations), and SXM (structured data movement) modules. Each LPU connects through 96 chip-to-chip links at 112 Gbps, for about 2.5 TB/s aggregate bi-directional IO, with NVIDIA pairing language around 150 TB/s of on-chip memory bandwidth per LPU.

LayerMemory / compute shapeWhat it is optimized for
Groq 3 LPU chip500 MB SRAM; ~150 TB/s on-chip BW; MXM/VXM/SXMDeterministic per-token FFN / MoE decode stages
LPX tray (8 LPUs)4 GB SRAM; up to 256 GB DRAM via fabric; 9.6 PFLOPS FP8Local assembly-line partition + host attach
LPX rack (256 LPUs)128 GB total SRAM; 40 PB/s; 315 PFLOPS FP8Rack-scale interactive decode next to Vera Rubin NVL72
Vera Rubin GPU (contrast)Large HBM capacity; high FLOPS; flexible schedulingPrefill, decode attention, training, high-concurrency throughput

Attention-FFN Disaggregation: how LPX and Rubin share a token

Interactive inference is not one kernel. Prefill builds the KV cache over a large context. Decode then walks a per-token loop where attention over that cache and feed-forward (or MoE expert) layers stress different bottlenecks.

NVIDIA's heterogeneous design, often called Attention-FFN Disaggregation (AFD), assigns the legs of that relay race to different engines:

  • Vera Rubin NVL72 GPUs: prefill, long-context processing, and decode attention over the KV cache.
  • Groq 3 LPX: latency-sensitive FFN and MoE expert execution inside decode.
  • Intermediate activations exchange each token so each engine runs the stage it is built for.

Here's why that matters. Throughput-optimized GPU racks can still look slow to a single agent user if per-token FFN work jitters under small batches. LPU silicon is built to keep that stage cycle-exact. GPUs keep the memory-heavy attention path where HBM capacity wins. You get a two-engine serving path instead of forcing one architecture to own both regimes.

NVIDIA Dynamo is the orchestration layer NVIDIA documents for making that split operational: classify requests, route prefill to GPU workers, run the AFD loop during decode, move interim tensors with low overhead, and keep tail latency stable under bursty traffic. Speculative decoding is a related pattern NVIDIA describes, with LPX generating draft tokens quickly while Rubin GPUs verify.

Production status, benchmarks, and early cloud adopters

On August 24, 2026, NVIDIA announced that Groq 3 LPX was in full production as an extension of the Vera Rubin platform for ultrafast token generation in agentic systems. The same release cites Artificial Analysis running a 100,000-token context benchmark on Gemma 4 31B, measuring about 3,400 output tokens per second on Groq 3 LPX (NVIDIA's companion technical blog rounds a related figure to 3,431 tok/s). NVIDIA frames that as world-class interactivity for that model and context length.

Nebius is named as the first AI cloud bringing LPX into production via Nebius Token Factory. NVIDIA says purpose-built inference cloud Groq plans to be among the earliest adopters. Those are go-to-market signals, not a guarantee that every hyperscaler SKU list will show LPX tomorrow.

On the efficiency slide, NVIDIA's architecture blogs claim that pairing Vera Rubin NVL72 with LPX can deliver up to 35x higher inference throughput per megawatt versus previous-generation GB200 NVL72 at high-interactivity, long-context operating points for trillion-parameter-class models, and frame up to 10x more revenue opportunity for premium interactive tiers. Those multipliers are vendor Pareto-frontier claims at specific operating points. They are not a universal conversion factor for every model size and batch shape.

💡
Inside Deep Tech's take: Buy LPX for interactive decode economics, not for a single FLOPS spreadsheet cell. If your agents need stable tokens-per-second per user at long context, the SRAM-first deterministic path is the story. If your invoice is dominated by large-batch prefill, training, or CUDA library lock-in, Vera Rubin GPUs (or commodity GPUs) still own the job.

Determinism as a power feature, not only a latency feature

NVIDIA's September 15, 2026 technical blog on deterministic execution explains why cycle-exact schedules matter beyond token rate. Once the compiler knows current demand per cycle across all 256 LPUs, it can drive Preemptive Power (PEP) and Clock Period Synthesis (CPS) to shape voltage and clock edges before spikes arrive.

NVIDIA reports internal testing with more than 60% less voltage droop and estimates a high single-digit percentage reduction in baseline voltage (and a low-double-digit percentage power reduction framing) for the same workload because power scales with the square of voltage. LPX rack power management sits alongside factory-level DSX MaxLPS and NVL72 Intelligent Power Smoothing. The shared goal is more useful tokens inside a fixed megawatt envelope.

How LPUs compare with GPUs, Cerebras, and hyperscaler ASICs

All of these are non-GPU AI silicon stories. They are not substitutes.

  • GPUs: still the default for training, mixed workloads, CUDA ecosystems, and high-concurrency batched serving.
  • Groq LPU / LPX: inference and interactive decode; SRAM + determinism; rack-scale with Rubin for AFD.
  • Cerebras WSE: wafer-scale training and inference with huge on-wafer SRAM and MemoryX/SwarmX, different access model.
  • Google TPU, AWS Trainium / Inferentia, Microsoft Maia: hyperscaler-first ASICs inside one cloud's SKU and software path.

Choose LPX when per-user interactivity and decode-stage latency are the product. Choose GPUs when portability and ecosystem gravity dominate. Choose a hyperscaler ASIC when you already live in that cloud. Choose Cerebras when wafer-scale single-device programming is the bet. Do not collapse those into one "AI accelerator" RFP line.

Cerebras Wafer-Scale Engine: A Full Guide to WSE-3 and CS-3
Companion non-GPU silicon guide: wafer-scale WSE-3 / CS-3, MemoryX/SwarmX, CSoft/PyTorch, and honest limits versus GPU clusters and hyperscaler ASICs.

Honest limits: when LPUs and LPX are the wrong tool

Name the downside before the purchase order.

  • SRAM capacity vs HBM models: 500 MB per Groq 3 LPU and 128 GB total rack SRAM are tiny next to hundreds of GB of HBM on a modern GPU. Large models require partitioning across many chips.
  • Prefill weakness relative to GPUs: in the AFD design, long-context prefill and decode attention stay on Rubin. LPU is not sold as the best prefill engine.
  • Training unsupported as the LPX product pitch: this is interactive inference / decode acceleration, not a CUDA training replacement.
  • Rack-level model capacity is a system problem: DRAM via trays helps, but you are buying a coordinated rack plus software, not a drop-in PCIe card with a giant HBM stack.
  • Compiler and determinism tradeoffs: static schedules buy latency and power predictability; they cost dynamic flexibility and depend on compiler maturity for your model family.
  • Licensing / platform nuance: NVIDIA uses Groq and LPU under license from Groq, Inc. GroqCloud and NVIDIA LPX are related architectures with different commercial paths. Do not assume identical APIs, SLAs, or SKUs.
  • When GPUs still win: CUDA libraries, multi-cloud rental, mixed train-and-serve fleets, high-concurrency throughput-first free tiers, and any workload that does not need extreme tokens-per-second per user.

If those limits dominate your environment, keep decode on GPUs (or on a hyperscaler inference ASIC you already operate) and revisit LPX when interactive agent loops show up as a first-class latency SLA.


FAQ

What is a Groq Language Processing Unit (LPU)?

An LPU is Groq's inference-focused processor built around software-first compilation, a programmable assembly-line architecture, deterministic execution, and on-chip SRAM instead of HBM-first GPU design. It targets fast, predictable LLM token generation rather than general-purpose training.

What is NVIDIA Groq 3 LPX?

Groq 3 LPX is NVIDIA's rack-scale low-latency inference accelerator for the Vera Rubin platform. It pairs 256 Groq 3 LPU chips with Vera Rubin NVL72 GPUs so GPUs handle prefill and attention while LPX accelerates latency-sensitive FFN and MoE decode work.

What are the key Groq 3 LPX rack specs?

Per NVIDIA's March 16, 2026 architecture blog: 315 PFLOPS FP8, 128 GB total SRAM, 40 PB/s on-chip SRAM bandwidth, 256 chips, and 640 TB/s scale-up bandwidth across 32 liquid-cooled trays.

How does Attention-FFN Disaggregation work?

AFD splits decode into engines: Rubin GPUs run attention over the KV cache, LPX runs FFN/MoE stages, and intermediate activations exchange each token. NVIDIA Dynamo orchestrates routing and keeps the serving path coherent under variable traffic.

Is Groq 3 LPX the same as GroqCloud?

No. GroqCloud is Groq's inference cloud powered by LPU technology. NVIDIA Groq 3 LPX is Vera Rubin platform silicon that uses Groq / LPU branding under license from Groq, Inc. Related architecture, different commercial and deployment paths.

When should teams stay on NVIDIA GPUs instead of LPX?

Stay on GPUs when training, CUDA libraries, multi-cloud portability, large HBM capacity, or high-concurrency batched throughput dominate. Evaluate LPX when per-user interactivity and stable decode latency at long context are the product requirement.

Are the 35x throughput-per-MW and 3,400 tok/s figures independent audits?

No. The ~3,400 tok/s Gemma 4 31B @ 100K context figure is Artificial Analysis benchmarking cited by NVIDIA. The up-to-35x throughput-per-MW and up-to-10x revenue opportunity claims are NVIDIA architecture comparisons at specific high-interactivity operating points. Measure your own models.

Does LPX replace HBM?

No. LPUs keep the hot working set in on-chip SRAM for bandwidth and determinism. GPUs still use large HBM for capacity-heavy phases such as long-context attention. Heterogeneous racks need both memory stories.


Monday's move is simple. If agent loops or interactive coding assistants are already on the latency invoice, pick one production decode path, ask whether AFD-style GPU+LPU serving (or GroqCloud LPU inference) beats your current GPU-only baseline on tokens per second per user and tail latency, and only then debate rack BOM. Datasheet PB/s can wait until that experiment finishes.