Vera Rubin vs Blackwell: how big is Nvidia's new leap?
Remember when every AI lab on the planet was begging for Blackwell? Waiting lists, rationing, the works. Well, Blackwell is about to become last year's model. Nvidia's next thing is called Vera Rubin, and it is rolling into the big cloud providers in the second half of 2026.
First, a small surprise: Vera Rubin is not really "a chip". It is a whole rack. Nvidia built six different chips to work as one machine: a CPU, the GPU, the switches that tie GPUs together, the network cards, a data processing unit and an Ethernet switch. Seventy two GPUs in one rack, all designed together from day one. You do not buy a part. You buy the whole beast.
The numbers, side by side
- Inference, the part where a model answers your questions: about 5 times faster than Blackwell GB200.
- Training: about 3.5 times faster.
- Memory: up to 288 GB of the new HBM4 per GPU, at up to 22 terabytes per second.
- Transistors: 336 billion. Blackwell had 208 billion.
- And the headline Nvidia loves most: up to 10 times lower cost per token.
That last one is the number to watch. Running these models is brutally expensive, and every lab is burning money on it. A tenfold drop in what each token costs to produce is the kind of thing that turns money-losing features into normal ones.
Hang on. Five times faster than what, exactly?
Good catch, because this matters. That 5x is measured against the original Blackwell GB200. Nvidia already shipped a beefed up version, Blackwell Ultra, and against that one the jump in inference is closer to 3.3 times. Still huge. But the vendor's favourite comparison is always the one that makes the new thing look biggest. Keep that in mind every time a launch slide says "5x".
And these are still promises. Real numbers come when the racks are running real workloads in real data centres. Until then, Vera Rubin is the new king on paper. Very likely the new king in practice too, just maybe a slightly smaller one than the slides suggest.