Level 1 - Absolute Beginner
Mistral is a computer company in France. It builds artificial intelligence, or AI. AI can read, write and answer questions.
On October 6 the company showed a new model. Its name is Mistral Large 4. People also call it Le Chonk, because it is very big.
The model can read about one million words of text at one time. That is like reading many long books together.
At the end of October the company will give the model away for free. Then other people can run it on their own computers.
- model
- A computer program that has learned from a lot of data.
- artificial intelligence
- Computers that can learn and answer questions.
- company
- A business that makes or sells things.
- free
- Costing no money.
- computer
- A machine that stores and works with information.
- text
- Written words.
- big
- Large in size or amount.
- chip
- A small electronic part inside a computer.
Level 2 - Elementary
On October 6 the French company Mistral AI released a public preview of its biggest model so far. The official name is Mistral Large 4, but the company and its users also call it Le Chonk, a joke about how large it is.
The numbers are unusual. The model has 1.05 trillion parameters in total, but it only uses about 49 billion of them for each word it processes. This design is called a mixture of experts, and it keeps the cost of running the model much lower than its size suggests. It can also look at pictures, and it can hold one million tokens of text in its memory at once.
Mistral says the model was trained from the beginning on 3,800 Nvidia Grace Blackwell graphics chips inside its own data centres in Europe. The training data covers more than 160 languages, including every official language of the European Union.
The model is available now through the company's paid service, at 1.36 dollars for a million words of input and 4.18 dollars for a million words of output. The company has promised to publish the weights, the actual trained numbers inside the model, at the end of October. It has not yet said which licence it will use.
- preview
- An early version shown before the full release.
- parameter
- One of the adjustable numbers a model learns during training.
- trillion
- One million million.
- token
- A small piece of text, usually a word or part of a word.
- data centre
- A building full of computers that run online services.
- weights
- The trained numbers that make up a model and allow others to run it.
- licence
- The legal rules saying how something may be used.
- input
- The information you give to a program.
Level 3 - Intermediate
Europe's best funded artificial intelligence start up has produced something its American and Chinese rivals mostly keep behind closed doors: a frontier scale model it intends to give away. On October 6 Mistral AI opened a public preview of Mistral Large 4, a model carrying 1.05 trillion parameters in total but activating only about 49 billion of them, roughly 4.7 percent, for any given token. Inside the company and among early users it has already acquired a nickname, Le Chonk.
The architecture is a granular mixture of experts, a design that routes each token to a small subset of specialised sub networks rather than running the whole model every time. The practical consequence is that a model with a trillion parameters can be served at something close to the cost of a fifty billion parameter one. A 1.6 billion parameter vision encoder handles image input natively, the context window stretches to one million tokens, and the system is presented as a hybrid that can either answer directly or reason at length before replying.
Provenance is part of the pitch. Mistral says the model was trained from scratch on 3,800 Nvidia Grace Blackwell accelerators inside its own European data centres, on material spanning more than 160 languages including every official European Union language, a claim aimed squarely at customers who worry about where their data is processed. The company has not yet published the expert count, the routing scheme or the layer layout, saying those details will arrive with the weights at the end of October. No licence has been named, which leaves the most important question about an open weight release unanswered.
Early numbers are promising but self reported. Mistral claims 93 percent on the Cybench cybersecurity benchmark and 82 percent on CyberGym end to end, where several closed frontier models score close to zero because they refuse the task outright, and 61.7 percent on the DeepSWE agentic coding test. A blind human evaluation run by Surge AI placed it second of five models with 3.74 out of 5, ahead of GLM-5.3 and Kimi K3 but behind Claude Opus 5 on 4.22. Pricing through the company's own interface is 1.36 dollars per million input tokens and 4.18 dollars per million output tokens, with cached input at 14 cents.
- frontier scale
- At the leading edge of what current systems can do.
- mixture of experts
- A design that sends each piece of input to a few specialised sub networks.
- route
- To direct something along a particular path.
- context window
- The amount of text a model can consider at one time.
- provenance
- The record of where something came from.
- benchmark
- A standard test used to compare systems.
- self reported
- Published by the party being measured rather than an independent tester.
Level 4 - Advanced
The competitive logic of frontier artificial intelligence has, until now, pointed in one direction: the larger the model, the tighter the secrecy. Mistral AI's public preview of Mistral Large 4, opened on October 6, inverts that instinct. The model carries 1.05 trillion total parameters while activating roughly 49 billion per token, about 4.7 percent of the whole, and the company has committed to publishing the weights at the end of October, which would make it comfortably the largest open weight system to emerge from Europe. The internal nickname, Le Chonk, has done more for the launch than any benchmark table.
Technically, the release is an argument about serving economics rather than raw capability. A granular mixture of experts routes each token to a narrow subset of specialised sub networks, so inference cost tracks the active parameter count rather than the headline figure, which is how a trillion parameter model reaches an interface price of 1.36 dollars per million input tokens and 4.18 dollars per million output tokens, with cached input at 14 cents. A 1.6 billion parameter vision encoder provides native image input, the context window extends to one million tokens, and the model is positioned as a hybrid capable of either immediate response or extended deliberation. The expert count, top k routing and layer layout remain undisclosed pending the weight release, as does the licence, which is the single variable that will determine whether the word open means anything in practice.
Sovereignty runs underneath all of it. Mistral states that the model was trained from scratch on 3,800 Nvidia Grace Blackwell accelerators in its own European facilities, on corpora covering more than 160 languages and every official European Union language. For regulated buyers in Europe, a competent model whose training location and weights are both knowable addresses a procurement problem that no amount of American capability solves, and it is a reasonable guess that this, rather than any leaderboard position, is the commercial thesis.
The performance claims deserve the usual caution, since most were produced in private evaluation. Mistral reports 93 percent on Cybench and 82 percent on CyberGym end to end, a domain where several closed frontier models register near zero because their safety training declines the task, alongside 61.7 percent on DeepSWE version 1.1, 59.4 percent on SWE-Atlas question answering and 28.3 percent on Terminal-Bench 4.0, for a combined coding agent index of 49.8 percent. An independent blind human evaluation conducted by Surge AI is the more useful signal: 3.74 out of 5, second of five systems, ahead of GLM-5.3 at 3.60 and Kimi K3 at 3.59, and behind Claude Opus 5 at 4.22. Safety testing shows resistance to 93.3 percent of attacks on Lakera's benchmark. Against DeepSeek V4 Pro, which already ships 1.6 trillion parameters under an MIT licence, Mistral's advantage is not scale but jurisdiction.
- inference
- The act of running a trained model to produce an answer.