Same price per token, so why is the bill so different?

Two models with an identical price list can end up with wildly different invoices. The trick is in what happens between the tokens.

Here is the price list. Sonnet 5.5, GPT-6 Sol and GPT-6.1 Sol all cost $2 per million input tokens and $10 per million output tokens. Astra, the OpenAI flagship, is $10 and $50. Looks like a tie.

It is not. A token price only tells you what a word costs, not how many words the model burns to think its way to an answer. One review put it neatly: Sonnet 5.5 is cheap per token but can be expensive per task. It reportedly used record amounts of tokens on the independent index. In one AutomationBench example, Sonnet cost about $1.14 per task while the new Sol cost about $0.30. In another independent measurement, Sonnet took roughly 859 seconds per task, against 392 seconds for GPT-6 Sol.

So is the price list lying to us? Not really. It is just answering a different question.

It tells you what a token costs. You care what a finished job costs. Anthropic claims Sonnet 5.5 is up to 30% cheaper per task than its predecessor, and that may well be true, but that is a comparison against its own last model, not against a rival.

So before you commit, run your own workload and look at the final invoice, not the price per million.