Ten times cheaper per token: what does that really mean?

It does not mean AI uses fewer tokens. It means something more interesting.

"Ten times lower cost per token." It is the line Nvidia keeps repeating about its new Vera Rubin hardware, and it sounds like it means AI is about to get ten times cheaper for everyone. Not quite. Let's unpack it.

First: which part of AI are we even talking about?

A model lives two lives. Training is when it gets built, a gigantic one-off job. Inference is everything after: every question you type, every answer it spits out, billions of times a day, forever. The tenfold figure is about inference. Training improves too, but by a smaller amount, roughly 3.5 times faster.

And inference is exactly where the money goes. You train a model once. You run it all day, every day, for millions of people. That is where labs are burning cash right now.

Second: cheaper tokens, not fewer tokens

Here is the bit people get wrong. New hardware does not make a model use fewer tokens. Ask the same question, get the same answer, same number of tokens. That depends on the model, not the silicon. What changes is how much it costs the lab to produce each one. Same token. A tenth of the price to make.

So will my AI subscription get ten times cheaper?

Don't hold your breath. A cheaper token is a saving for the lab, and every lab gets to decide what to do with it. Some might cut prices to win customers. Some might just keep the difference and finally fix their margins. And some will do the sneaky clever thing: keep the price the same and let the model think ten times longer, run agents for hours, chew through huge documents. Same bill for you, much more work behind it.

So the honest answer is: it depends who you are buying from. The hardware hands labs a pile of savings. What reaches you is a business decision, not a law of physics.