Is new hardware the only way to make AI cheaper?

Nvidia sells the shiny half of the story. The other half is a trick called caching.

When Nvidia launches a new monster chip and promises tokens ten times cheaper, it is easy to think that is how AI gets cheaper: better silicon, end of story. It is only half of it. There is a second, much quieter way labs are cutting costs, and it has nothing to do with buying hardware.

Saving number one: make each token cheaper

This is the hardware route. Nothing about your request changes. The model thinks just as hard, digs just as deep, writes just as many tokens. Each token simply costs less electricity and less machine time to make, because the chip is better. Pure engineering.

Saving number two: stop doing the same work twice

Now the clever bit. A huge amount of what gets sent to an AI model is repeated. The same long set of instructions at the start of every request. The same document pasted into a conversation that keeps going. Why compute all that from scratch every time?

So labs stopped doing it. It is called prompt caching. If the start of your request matches something the model already processed recently, it reuses that work and only computes what is new. On that repeated chunk, the cost can drop by around ninety percent.

Chinese labs, DeepSeek most of all, made cheap cache hits a headline part of their pricing and leaned into it hard. Now it is standard at the big Western labs too. Add to that routing, where easy questions quietly go to a small cheap model and only the hard ones reach the expensive one, and you have a whole software toolkit for saving money.

Wait, so does that make the hardware upgrade less important?

Nope, and this is the fun part. The two do not add up, they multiply. Hardware makes every token cheaper. Caching means far fewer tokens need computing from scratch in the first place. A lab with both gets savings stacked on savings. Which is why that famous "10x" from Nvidia is only part of the picture. The labs that win on cost will be the ones that pull both levers at once.