HBM3E vs HBM4: Specs, Vendors and GPUs (2026)

Back
Team Aquanode

Team Aquanode

Sarthak Vaish

Updated OCTOBER 8, 2026Published OCTOBER 8, 2026

HBM4 is the stacked-memory standard that follows HBM3E: it doubles the interface to 2,048 bits per stack, which JEDEC says lifts a stack to up to 2 TB/s at 8 Gb/s per pin. HBM3E is what today's rentable GPUs use, including the H200, B200, B300 and MI355X, while HBM4 is arriving with NVIDIA Vera Rubin and AMD's MI400 generation, as of October 2026.

This guide covers:

  • What changed between HBM3E and HBM4, from the JEDEC standard
  • What SK hynix, Micron and Samsung have published about their HBM4 parts
  • Which GPUs use which memory, with vendor-published capacity and bandwidth
  • Why memory bandwidth matters more than capacity for inference, and when it does not

TL;DR

  • HBM4 doubles the I/O count per stack (2,048 vs 1,024 bits) and the channel count (32 vs 16). JEDEC's standard specifies up to 8 Gb/s per pin and up to 2 TB/s per stack. Vendors say their parts run faster than that.
  • Today's rentable GPUs use HBM3E: 141 GB at 4.8 TB/s on the H200, and 288 GB at 8 TB/s on the B300 and on AMD's MI355X.
  • HBM4 GPUs: NVIDIA's Vera Rubin lists 288 GB of HBM4 at 19.2 TB/s per GPU. NVIDIA's page calls the product "Available Now" and "ramping into full production." NVIDIA's HGX page and Rubin technical blog list 22 TB/s for the same GPU, so the two NVIDIA pages disagree.
  • Verdict: for inference of large models, memory bandwidth sets your tokens per second, and HBM4 is a large step. But you rent what exists. Today that means HBM3E parts, and the 288 GB ones are the capacity leaders.

What HBM is, in one paragraph

High Bandwidth Memory stacks DRAM dies vertically and puts the stacks beside the GPU on the same package, connected by a very wide bus. Width is the trick: instead of a few fast pins, HBM uses thousands of slower ones. Our glossary has the basics at HBM and VRAM. The rest of this post is about the generation change.

HBM3E vs HBM4 at the standard level

JEDEC published the HBM4 standard (JESD270-4) in April 2025. Per JEDEC's announcement:

HBM3E (as shipped)HBM4 (JEDEC JESD270-4)
Interface width per stack1,024 bits2,048 bits
Independent channels per stack16 (HBM3 generation)32, each with 2 pseudo-channels
Pin speed in the standardNot covered by the HBM4 standardUp to 8 Gb/s
Bandwidth per stackSee note belowUp to 2 TB/s
Stack heightsUp to 12-high on today's parts4, 8, 12 or 16-high
Die densityNot covered here24 Gb or 32 Gb
Max capacity per stack36 GB on MI355X (per AMD docs)64 GB (32 Gb, 16-high)
Voltage optionsNot covered hereVDDQ 0.7, 0.75, 0.8 or 0.9 V; VDDC 1.0 or 1.05 V

A note on HBM3E bandwidth per stack. JEDEC's HBM4 announcement does not restate it, and we did not find a vendor page that gives a per-stack HBM3E figure we could cite. We can compute one from AMD's published figures: AMD's ROCm documentation describes MI355X's HBM3E at 8 Gb/s per pin with 36 GB per stack and 8 TB/s in total. Eight stacks at 1,024 bits and 8 Gb/s gives 1 TB/s per stack, which matches. That is our arithmetic on AMD's numbers, not a vendor-stated per-stack figure, and it applies to that one product.

JEDEC also says the HBM4 interface is backwards compatible with HBM3 controllers, so one controller can serve both.

The doubled width is the headline. Even at the same pin speed as a good HBM3E part, a 2,048-bit stack moves twice the data per cycle. And because vendors push pin speeds past the standard's 8 Gb/s, the real per-stack number is higher than 2 TB/s, as the next section shows.

What the memory makers say

All three DRAM makers have announced HBM4. The figures below are vendor claims.

SK hynix. Its September 2025 announcement said HBM4 development was complete and the part was ready for mass production, with operating speed above 10 Gb/s ("far exceeded" the JEDEC 8 Gb/s), 2,048 I/O terminals, and power efficiency improved by more than 40% over the previous generation. It uses 1b-nm (fifth-generation 10 nm-class) DRAM and the MR-MUF packaging process. The company projected up to 69% better AI service performance, a company projection and not a measured result. On its Q2 2026 earnings call, the company said it began HBM4 mass production and supply to key customers in the second quarter and will expand production in the second half, with HBM4E targeted for full-scale mass production in 2027, per press coverage and the call transcript.

Micron. Micron said it began volume shipment of HBM4 36GB 12-high in the first quarter of 2026, designed for NVIDIA Vera Rubin. Its claims: over 11 Gb/s pin speed, over 2.8 TB/s per stack, about 2.3 times the bandwidth of HBM3E, and more than 20% better power efficiency than HBM3E. It has also sampled a 16-high 48 GB stack, which is 33% more capacity per placement than the 36 GB part. Check: 11 Gb/s times 2,048 bits divided by 8 is about 2.8 TB/s, so the claim is internally consistent.

Samsung. Samsung's February 12, 2026 release says it began mass production of HBM4 and shipped commercial products to customers, with a consistent speed of 11.7 Gb/s that can be pushed up to 13 Gb/s, up to 3.3 TB/s per stack, and 24 GB to 36 GB capacities in 12-high stacks. The release does not name a customer; press coverage links the part to NVIDIA's Rubin. At 13 Gb/s and 2,048 bits the arithmetic gives about 3.3 TB/s per stack, matching Samsung's figure. These are Samsung's claims, not independent measurements.

One useful pattern: reports say NVIDIA pushed for pin speeds above 11 Gb/s, which forced redesigns at all three vendors and delayed mass production. That is a reminder that the JEDEC number (8 Gb/s) is a floor for the market, not what flagship GPUs ask for.

Which GPUs use which memory

These are the vendor-published figures. Where a vendor page does not give a number, we write "not published on that page."

GPUMemoryCapacityBandwidthSource
NVIDIA H100 SXMHBM3 family (type not stated on the product page)80 GB3.35 TB/sNVIDIA H100 page
NVIDIA H200HBM3e141 GB4.8 TB/sNVIDIA H200 page
NVIDIA B200HBM3e192 GB (implied by NVIDIA's "50% more than Blackwell" for B300)Not published on the pages we readNVIDIA developer blog
NVIDIA B300HBM3e288 GB8 TB/sNVIDIA developer blog
AMD MI355XHBM3E288 GB8 TB/sAMD product page
NVIDIA Vera RubinHBM4288 GB19.2 TB/sNVIDIA Vera Rubin NVL72 page

Reading across the table:

  • H100 to H200 was a memory-only upgrade in the same Hopper generation: 80 GB to 141 GB, 3.35 TB/s to 4.8 TB/s. NVIDIA's H200 page calls this 1.4 times the bandwidth.
  • B300 vs H100. NVIDIA's blog describes 288 GB of HBM3e as 3.6 times the on-package memory of H100 and 8 TB/s as a 2.4 times improvement over H100's 3.35 TB/s.
  • MI355X vs B300 have the same capacity and the same headline bandwidth. Their differences are elsewhere.
  • Rubin vs B300. NVIDIA keeps 288 GB per GPU, and lists 19.2 TB/s of HBM4 bandwidth against 8 TB/s for B300. That ratio is 2.4 times, computed from the two vendor figures. NVIDIA's HGX page and technical blog give 22 TB/s for the same GPU (about 2.75 times B300, computed), so the ratio depends on which NVIDIA page you read.
  • AMD's next step. AMD's Helios rack page lists 31 TB of HBM4 across a 72-GPU rack. We cover the GPU in our MI400 and MI450 guide.

For the Rubin chip itself, see our Rubin guide. For the Blackwell Ultra part, see the B300 guide.

Why bandwidth matters for inference

When a model generates one token at a time, each step reads the active weights and the KV cache from memory. The arithmetic is light compared with the data moved, so the GPU waits on memory. Two consequences:

  • A rough ceiling on single-stream speed. If a step must read a given number of gigabytes, the memory bandwidth divided by that size bounds tokens per second. This is a bound, not a prediction, and batching, caching and quantization move it. We do not publish a tokens-per-second figure for any card here, because we have no cited throughput to base one on.
  • Capacity decides whether it fits at all. A 288 GB GPU can hold weights and KV cache that would force a split across two 141 GB GPUs. Splitting across GPUs brings in NVLink traffic, which is its own cost.

Training leans on bandwidth too, but compute and interconnect usually share the bottleneck with it. HBM4's doubled width is most visible on memory-bound work: long-context decoding, large batch decode of big mixture-of-experts models, and anything with a large KV cache.

Power, packaging and supply

Wider stacks are not free. SK hynix claims a 40% power-efficiency gain over its previous generation and Micron claims 20% over HBM3E, which matters because memory is a meaningful part of a GPU's power budget. HBM4 also changes how the stack is built: the base die can be fabricated on a logic process (SK hynix says its base die is made in collaboration with TSMC), which is why HBM4 is more tightly coupled to each GPU design than earlier generations.

Supply is the practical limit. Coverage from mid-2026 describes Micron's 2026 HBM4 output as already committed under contracts, and SK hynix said on its Q2 call that it would expand production in the second half. In plain terms: HBM4 GPUs will be scarce while HBM3E parts remain the volume product.

Cost

Memory is a large share of a high-end GPU's cost, and a GPU with more or faster HBM rents for more per hour. We do not type a rental price here. The live from-prices for three HBM3E parts are in the box below. The useful comparison is price per GB of memory and per TB/s, which you can compute from the vendor figures in the table above and the live price.

Rent today

Three HBM3E GPUs with very different memory sizes: the H200 (141 GB), the B300 (288 GB) and the MI355X (288 GB).

What's next

HBM4E is the next step. SK hynix said on its Q2 2026 call that HBM4E is proceeding on schedule, targeting full-scale mass production in 2027. NVIDIA lists Vera Rubin at 288 GB of HBM4 per GPU, and Micron has sampled 48 GB 16-high stacks that could raise capacity per placement by a third. We do not state product dates beyond what these sources say.

For every datacenter chip on one page, with status, memory and links to each guide, see the datacenter GPU guide.

FAQ

What is the difference between HBM3E and HBM4?

HBM4 doubles the interface from 1,024 to 2,048 bits per stack and the channels from 16 to 32. JEDEC's standard gives up to 8 Gb/s per pin and up to 2 TB/s per stack, with up to 16-high stacks and 64 GB per stack.

Which GPUs use HBM4?

NVIDIA's Vera Rubin lists 288 GB of HBM4. AMD's Helios rack lists 31 TB of HBM4 across 72 GPUs. Neither is the volume product you can rent today, which is HBM3E.

How fast is HBM4?

JEDEC's standard says up to 2 TB/s per stack. Micron claims over 2.8 TB/s per stack at over 11 Gb/s, and Samsung claims its part at up to 13 Gb/s and 3.3 TB/s per stack.

Is more HBM capacity or more bandwidth better for LLM inference?

Capacity decides whether a model and its KV cache fit. Bandwidth decides how fast tokens come out once they fit. A larger model needs capacity first.

Is HBM3E obsolete?

No. The H200, B200, B300 and MI355X all use it, and those are the GPUs available now.

Sources

#datacenter gpu#gpu interconnect#hbm4#hbm3e#gpu memory#memory bandwidth

Submit the job. Everything after that is ours.

Sign up in 60 seconds. Pay for the GPU minutes you actually use.

© 2026 Aquanode. All rights reserved.

All trademarks, logos and brand names are the property of their respective owners.