What is actually inside Mistral's Le Chonk?
Start with the trick. Le Chonk is a mixture-of-experts model. Picture a giant team of specialists where, for every chunk of text, only a small group clocks in. In total it has roughly 1.05 trillion parameters, but only about 49 to 52 billion are active at any one time. That is why running it does not bankrupt you despite the headline size.
It takes in text and images and answers in text. It was trained from scratch on about 3,800 Nvidia Grace Blackwell chips, in Mistral's own European data centres. The public preview landed on 6 October.
Ok, and how much can it read in one go? Well, that depends on who you ask.
Mistral's own model card says a context window of one million tokens. The independent tracker Artificial Analysis measured roughly 524,000. Another benchmarking site lists 512,000, with output capped at 256,000. Nobody has publicly explained the gap yet.
Until someone does, the sensible move is to plan around the lower number. If the full million shows up later, that is a bonus, not a foundation to build on.