
Mistral Large 4 Has 1.05T Parameters—but Its Scores Need Verification

On October 6, 2026, Mistral AI put Mistral Large 4 into public preview: its model card lists 1.05 trillion total parameters, 52 billion active parameters, a one-million-token context window, and current API rates of $0.68 per million input tokens and $2.09 per million output tokens. The available product is hosted access; downloadable weights are planned for a later release.
In its October 6 launch post, Mistral instead gives 49 billion active parameters, offers the preview API through Mistral Studio, reports 61.7% on DeepSWE v1.1, and promises weights by the end of October. In Le Monde’s report, co-founder Guillaume Lample said the model “narrows the gap” with leading systems; the paper carried a preliminary 63% Deep SWE 1.1 figure, noted that independent rankings had yet to confirm the performance, and gave October 27 as the planned open-weight date.
What the public preview gives developers
The preview makes Mistral Large 4 callable as a hosted model under the identifier mistral-large-4. It accepts text and image input and supports structured output, function calling and document question answering. Those capabilities describe the API that developers can use now. Running the model on private infrastructure depends on the future distribution of its weights, so the “Open” label attached to the preview should be read alongside the release timetable.
The one-million-token context window is a capacity figure, not a measured guarantee that every detail in a large repository or document collection will be retrieved accurately. A larger window permits longer requests, but the quality of answers to long inputs depends on the task and evaluation conditions. That distinction matters here because the launch pitches repository understanding and complex coding workflows alongside the unusually large context specification.
Mistral Large 4 claims, by status
- Total size — published specification: 1.05 trillion parameters describes the full mixture-of-experts model, including parts that are not active for every token.
- Active size — unresolved discrepancy: 52 billion and 49 billion are both published active-parameter figures. The difference has no published explanation.
- Access — available now: the public-preview API can be used through Mistral Studio. The model identifier is mistral-large-4; downloadable weights belong to the later release.
- Context — advertised capacity: one million tokens is the listed window, rather than an independent test of long-context recall or coding performance.
- Price — currently displayed rates: input costs $0.68, cached input $0.07 and output $2.09 per million tokens at the discounted prices shown on the model listing.
- DeepSWE — reported performance: 61.7% is the vendor’s published DeepSWE v1.1 result; a separate preliminary account gives 63%. Published details do not establish that the figures came from an identical run.
- Open weights — planned: the end-of-October release window includes the more specific October 27 date. The weights have not yet been released.
In a mixture-of-experts system, total and active parameters measure different things. The larger figure counts the whole model; the active figure represents the smaller part used when processing a token. The two active-parameter figures, by contrast, purport to describe the same property of the same model, which is why their difference is material even though it is small relative to the total.
The API price and the performance claim
The posted discount sits beside higher reference rates of $1.36 per million input tokens, $0.14 per million cached input tokens and $4.18 per million output tokens. The displayed discount halves each of those figures. As a purely illustrative calculation, a request billed for one million fresh input tokens and one million output tokens would cost $2.77 at the displayed discount, compared with $5.54 at the reference rates. Actual bills depend on how much input is cached and which tokens qualify for each category.
That distinction is relevant when comparing the preview with another hosted model. A workload that repeatedly sends the same long prompt may have a different effective price from one that sends only new input, even if the output volume is identical. The current rate card also describes API use; it cannot establish the operating cost of a future self-hosted deployment without details of the released weights, hardware and throughput.
The coding claim has a separate evidence problem. DeepSWE v1.1 concerns longer software-engineering tasks, but a benchmark percentage is meaningful only with its model version, tools, test setup and scoring rules attached. The published 61.7% and preliminary 63% are close, yet there is no basis in the available accounts to treat them as measurements from an identical run or as an independently confirmed result for that benchmark. The preview can be evaluated on a buyer’s own tasks today, but that is a different kind of evidence from a comparable third-party benchmark.
What changes when the weights arrive
Releasing weights would give open-model users access to the model files and allow deployment choices that the hosted preview cannot provide. It would also let outsiders inspect the distributed version and test whether its measured performance matches the preview’s published claims. The period before release includes red-teaming with partners and state authorities, with further architecture and evaluation details expected alongside the weights.
October 27 remains the more specific planned date, within the broader end-of-month pledge. Whether the distributed version retains the preview’s scores, and which active-parameter figure describes it, will become clearer when the weights and fuller test details are available.
Also read:
Related articles
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.


