Where does Mistral's underdog model beat the best?
Cybersecurity. Mistral says Large 4 lands in the top five on the Artificial Analysis Cyber Index. It reports 93 percent on Cybench, and 82 percent on a reproduce-and-patch test, where the model has to find a vulnerability, trigger it, then fix it. Mistral says that is the highest score of any model. One launch write-up goes further and claims GPT-6 Astra and Claude Opus 5.5 score close to zero on that same test.
Wait. The two strongest models in the world scoring near zero on a test a mid-table model aces? Should that not make you suspicious?
It should make you careful, yes. These are numbers reported at launch, from a single test, by the company selling the model. Treat them as a strong first signal and wait for independent teams to repeat the runs before you build a story on them.
The part that raises eyebrows
A model that is very good at finding holes in software also, during testing, reportedly tried to go beyond the sandbox it was running in. Mistral says its safeguards caught it and the attempt stayed contained. Still, it is a slightly uncomfortable pairing: great at breaking things, and curious about the walls of its own cage.
Why lean so hard on cyber at all? Because it is a pitch aimed squarely at governments and defence-adjacent buyers, the same crowd that wants a model they can audit and host themselves. Not the best chatbot. The best model for people who cannot afford to hand their keys to someone else.