Sonnet 5.5 or GPT-6.1 Sol: who actually wins?

Everybody wants a clean answer. The honest one is messier, and it starts with a comparison that is missing.

You would think this is easy. Two mid-tier models, one day apart. Just line up the scores. Except most of the head to head charts compare Sonnet 5.5 with the older GPT-6 Sol, not the new one. Against that older model, Sonnet wins 18 of 19 shared benchmarks.

Against the new Sol, the evidence is patchy. On AutomationBench, a test of multi-step business workflows, Sonnet 5.5 scored 44.7%. The new Sol landed somewhere between about 32% and 36%, depending on the effort setting. On DeepSWE, a test built on real codebases, OpenAI reports 75.2% for the new Sol, ahead of the best Astra result. There is no confirmed Sonnet 5.5 number to set next to it yet.

Then there is the price of a win. In the AutomationBench run, Sonnet cost about $1.14 per task. The new Sol cost about $0.30.

So who wins? Hate to say it, but that is the wrong question.

The better question is what job you have. If you are chasing the highest score on broad, long knowledge work, Sonnet has the edge. If you are running thousands of agent tasks and every cent counts, the new Sol's price per task is hard to ignore.

And keep one thing in mind: most of the new Sol numbers come from OpenAI itself, and every company picks the charts that flatter it.