Can a mid-tier model really outrank the flagships?

One of the cheaper models just landed second on a major independent ranking, ahead of OpenAI's top model.

Meet Sonnet 5.5, Anthropic's mid-tier model, released on 28 September. On the Artificial Analysis Intelligence Index, a widely watched independent scoreboard, it landed in second place. Its score at max effort is 56. GPT-6 Sol got 48, and OpenAI's flagship Astra got 53.

It is not just one chart, either. In Vals AI's comparison against GPT-6 Sol, it wins 18 of 19 shared benchmarks. On Terminal-Bench 4.0, an agentic coding test, it scored 70.6%, up from 10.3% for its predecessor. Anthropic also claims it is more than 30% faster than Sonnet 5 and up to 30% cheaper per task. On a real-world work test, it sits about two points behind Anthropic's own flagship, Opus 5.5.

Wait. 10.3% to 70.6%? That is not an upgrade, that is a different animal. Can that be real?

One reviewer flagged exactly this: the old score looks unusually low, so treat the jump as a rough signal, not a precise one. And the independent ranking was run on a pre-release version with a bug affecting structured outputs. Anthropic says it fixed the bug for the public release and that the results might even be slightly understated.

Both caveats are worth knowing, and neither changes the big picture. The middle of the lineup keeps creeping up on the top, and this time it got ahead on a scoreboard people actually watch.