Mistral Large 4 : A Trillion-Parameter Preview With Open Weights on the Way

Mistral Large 4 brings 1.05T parameters, 49B active parameters, multimodal reasoning, and promised open weights—while benchmark, pricing, and context-window details still need scrutiny.

•
10 min read
Ankit Singh
Software Engineer
Mistral Large 4 : A Trillion-Parameter Preview With Open Weights on the Way

On October 6, 2026, Mistral launched a public preview of Mistral Large 4, its largest model to date and, by the company's description, its most capable. The preview API is available through Mistral Studio, while Mistral says the model weights will be released by the end of October.

The headline specifications are substantial: 1.05 trillion total parameters, 49 billion active parameters, and a 1.6 billion-parameter vision encoder. Large 4 is a multimodal hybrid instruct-and-reasoning model that accepts text and images as input. Mistral describes it as a granular Mixture-of-Experts model designed for coding, agentic workflows, cybersecurity, knowledge work, and multimodal tasks.

There is one important caveat up front: the model is not yet downloadable. Mistral calls Large 4 an open-weight model and promises the weights later this month, but as of October 6 the public preview is still API-only and the final licensing terms for the released weights have not been published. Artificial Analysis currently classifies the preview as proprietary because its weights are not yet available. Mistral Mistral Docs Artificial Analysis

Big on disk, lighter at runtime

Large 4 uses a granular Mixture-of-Experts (MoE) architecture. Instead of activating the whole network for every token, the model routes each token through only part of its available parameters.

That is why the 49B active-parameter figure matters. It is a useful first-order indication of the amount of model computation involved per token, while the full 1.05T parameter model still has to be stored across the inference system unless weights are aggressively quantized or otherwise compressed.

Only about 4.7% of the total parameters are active at a time:

49B / 1.05T ≈ 4.7%

That does not make Large 4 a 49B model, though. Weight memory, sharding, KV cache, runtime overhead, and the chosen precision all still matter when deploying it.

Here is how it compares with its predecessor:

Mistral Large 3

Mistral Large 4 (preview)

Released

December 2, 2025

October 6, 2026

Total parameters

675B

1.05T

Active parameters

41B

49B

Context window

256K

1M advertised by Mistral

Input

Text + image

Text + image

License

Apache 2.0

Not yet announced

Weights

Released

Promised by end of October

The jump is meaningful. Total parameters are up by roughly 56%, active parameters by about 20%, and Mistral's advertised context window is almost four times larger than Large 3's.

Mistral says Large 4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacenters. A significant share of the training data was multilingual, covering more than 160 languages, including every official language of the European Union. Mistral Large 3 Mistral Large 4

What Mistral claims

The launch announcement contains a long list of benchmark results. It is important to distinguish between Mistral's reported results and numbers that have already been reproduced independently.

Cybersecurity

Mistral says Large 4 ranks among the top five models globally on the Artificial Analysis Cyber Index and leads open-weight models developed outside China.

On one vulnerability-reproduction-and-patching test, Mistral reports 82%, while on Cybench it reports 93%.

Those results come with an unusual caveat. Mistral says several leading closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on the vulnerability test because they refuse to perform the requested security work.

That means the result measures more than raw technical capability. It also reflects each provider's safety policy. A model that refuses vulnerability reproduction may score worse on this particular benchmark even if the refusal is intentional and desirable from that provider's perspective.

Mistral is leaning into that distinction. It says vetted partners, cybersecurity leaders, and state authorities will receive a reduced-moderation version with expanded cyber capabilities for red-teaming before the public weights arrive.

Coding

Mistral reports:

  • 61.7% on DeepSWE v1.1

  • 59.4% on SWE-Atlas-QnA

  • 28.3% on Terminal-Bench 4

  • 49.8% on its combined Coding Agent Index

Mistral says the Coding Agent Index result places Large 4 ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max.

The important provenance detail is that Mistral explicitly says the DeepSWE, Terminal-Bench 4, and SWE-Atlas-QnA figures use numbers reported by Artificial Analysis. These were preliminary evaluations published as part of Mistral's launch material, rather than a fully public, independently reproducible ranking for the preview at launch time. Mistral Large 4

Mistral also ran a blind human evaluation with Surge AI. Professional annotators rated the outputs on a 1–5 scale with model identities hidden. Large 4 Preview scored 3.74, placing second among five evaluated models, behind Claude Opus 5 at 4.22 and ahead of Kimi K3, GLM-5.3, and GLM-5.2.

There is a broader reason not to interpret the coding numbers as proof that Large 4 is the overall coding leader: benchmark results can change considerably depending on the model version, agent harness, tools, and configuration. Current public coding leaderboards show stronger results from some competing systems under their best published configurations. VentureBeat

Agentic work

Mistral reports 59.9% on AutomationBench, covering 657 business workflows across applications such as Gmail, Google Sheets, Slack, and Salesforce.

It also reports 1,393 Elo on AA-Briefcase, Artificial Analysis's long-horizon knowledge-work evaluation.

These results are particularly relevant because Large 4 is positioned less as a chatbot and more as a general-purpose model for agents that gather information, use tools, and produce finished work.

Vision

On the Dense 200 visual-grounding benchmark, Mistral reports 42%, compared with 41% for GPT-6 Astra.

That is an impressive comparison, but the exact competitor configurations shown in Mistral's launch material are not all independently documented in the public benchmark sources. The safest interpretation is therefore that Mistral reports a 42% result that exceeds the GPT-6 Astra figure in its comparison, rather than treating this as a universally reproduced head-to-head result. Mistral Large 4 VentureBeat

The launch also includes third-party evaluations by Vals AI.

Mistral says Large 4 outperformed GPT-6 Astra on representative finance and legal tasks and beat all open-source models on Harvey's Legal Agent benchmark.

Vals currently reports:

  • 54.68% on Finance Agent v2

  • 15.83% on Harvey's Legal Agent Benchmark

  • 48.05% on the overall Vals Index

On Harvey's benchmark, that 15.83% result ranks 6th out of 75 models in the current Vals dataset.

That distinction matters. Saying Large 4 “beats all open models” does not mean the model solves 80% or 90% of the legal benchmark. Its absolute pass rate is only about 16%; its advantage is primarily relative to the other models tested on the same benchmark. Vals AI

The first independent read

The most useful independent snapshot so far comes from Artificial Analysis.

Artificial Analysis currently scores Mistral Large 4 Preview at about 38 on its Intelligence Index, substantially above Mistral Large 3 at 9 and Mistral Medium 3.5 at 14.

The number is a major improvement for Mistral, but it does not establish Large 4 as the strongest open-weight model overall.

In the current open-weight landscape, Large 4 would sit behind seven Chinese models in the ranking visible from current third-party comparisons. Because the preview weights are not yet public, Artificial Analysis currently labels Large 4 Preview as proprietary rather than open-weight.

So the precise claim is:

Large 4 is currently the strongest non-Chinese open-weight model in the relevant comparison, but it does not currently lead the global open-weight field.

That is a more defensible conclusion than simply calling it the world's best open model.

Artificial Analysis also highlights an interesting characteristic that can easily get missed in headline benchmark scores: Large 4 is unusually verbose.

In the Intelligence Index evaluation, it generated about 200 million output tokens, compared with a median of roughly 81 million among the comparison models.

That has practical consequences. A model can have a competitive per-token price and still become expensive if its reasoning and responses are consistently much longer than competing models. Artificial Analysis

Pricing: there is a discrepancy worth watching

Mistral's Large 4 documentation currently displays:

  • $0.68 per million input tokens

  • $0.07 per million cached input tokens

  • $2.09 per million output tokens

The same page also shows higher figures — $1.36 input and $4.18 output — crossed out beside the current lower rates.

Artificial Analysis and Vals, however, currently use the $1.36 / $4.18 figures in their launch-day model evaluations.

That means two different pricing references are circulating at the same time:

Source/context

Input

Cached input

Output

Current Mistral documentation

$0.68/M

$0.07/M

$2.09/M

Artificial Analysis / Vals evaluation pricing

$1.36/M

—

$4.18/M

For someone actually building against the API, the current Mistral documentation is the number that matters. For comparisons involving published third-party benchmark costs, the higher figures may still be used.

So benchmark cost figures should not automatically be interpreted as the same thing as today's live API tariff.

The low cached-input price is especially useful for workloads involving large prompts that are repeatedly reused, such as long system instructions, document collections, or multi-step agents. Mistral Docs Artificial Analysis Vals AI

The 1M context claim needs a footnote

Mistral advertises a 1 million-token context window for Large 4.

That is the number worth quoting when describing Mistral's own specification. But independent model trackers currently list lower figures:

  • Artificial Analysis: 524K

  • Vals AI: 512K

Artificial Analysis defines its context figure as the maximum combined input-and-output context, while Vals currently lists a 256K maximum output limit.

The reason for the difference between Mistral's advertised 1M figure and the third-party values is not yet clear.

So the safest wording is:

Mistral advertises a 1M-token context window, but current third-party listings record roughly 512K–524K. The discrepancy remains unresolved.

For developers building long-context applications, that is a meaningful detail. The practical limit can depend on the exact API endpoint, provider, maximum output setting, and model configuration, so a real workload should be tested rather than assuming every deployment exposes the full advertised window. Mistral Docs Artificial Analysis Vals AI

Self-hosting: this is still a very large model

The word “open-weight” can make a trillion-parameter model sound easier to run than it really is.

At 16-bit precision, storing 1.05T parameters requires roughly:

1.05T × 2 bytes ≈ 2.1 TB

At 8-bit precision, the raw weights are roughly:

1.05T × 1 byte ≈ 1.05 TB

Those are weight-storage estimates only. Real inference also needs memory for the KV cache, runtime overhead, CUDA or accelerator allocations, networking, and other serving infrastructure.

Large 4 therefore remains a serious multi-GPU deployment even though only 49B parameters are active for each token.

Mistral says the model will be deployable in private cloud and on-premise environments, and that it will be offered across several regions, including a European deployment operated end-to-end by Mistral under European law.

For API developers, the documentation lists support for structured outputs, function calling, document QnA, prefix completion, batching, and the Agents & Conversations APIs with built-in tools. Mistral Docs

Mistral is also positioning Large 4 around sovereignty

The hardware story is almost as important as the benchmark story.

Mistral says the model was trained and is being served on infrastructure in its own European datacenters. The company is explicitly framing this as an AI sovereignty play: organizations can use frontier-level capabilities while maintaining more control over where their models and data run.

That positioning is particularly relevant to industries with strict data-residency, auditability, or operational-control requirements.

It also explains why cybersecurity features get so much attention in the announcement. Mistral argues that a defender may need to reproduce a vulnerability to verify it, even when a hosted model refuses to perform that task.

The open-weight release is therefore not just a developer feature. It is a central part of Mistral's pitch around control, self-hosting, and independence from provider-level policy decisions. Mistral

What we still don't know

License and weights

This is the biggest unresolved issue.

Mistral says the weights will arrive by the end of October, but as of October 6 they are not publicly downloadable.

The final license also has not been announced. Large 3 used Apache 2.0, but there is no basis for assuming Large 4 will use the same terms.

For commercial users, this may matter more than a few benchmark points.

Architecture and post-training details

Mistral says more information on the model architecture and post-training methodology will be released alongside the weights.

That means important implementation questions remain open, including exactly how the routing architecture is structured, what the final checkpoint contains, and what inference optimizations will be supported in the public release.

The preview is not necessarily the final model

Mistral explicitly says the reinforcement-learning run behind the preview is still in flight and that the model continues to improve.

That means the checkpoint released later this month may not be identical to the public preview available today.

For developers evaluating Large 4 now, today's measurements should therefore be thought of as preview-era results, not necessarily final-release performance.

The takeaway

Mistral Large 4 is a major step forward for Mistral.

Its 1.05T total parameters and 49B active parameters make it one of the largest publicly announced MoE models, while the preview combines multimodal input, reasoning, coding, agentic workflows, and enterprise-oriented capabilities in a single system.

The independent evidence is promising. Artificial Analysis puts Large 4 far ahead of Mistral's previous Large 3 on its Intelligence Index, while Vals shows competitive results on finance and legal-agent tasks. Mistral's own evaluations are particularly strong in cybersecurity, where its willingness to perform vulnerability analysis gives it an advantage against systems with stricter refusal policies. Artificial Analysis Vals AI Mistral

But the launch does not establish Large 4 as the world's strongest open-weight model overall. Current independent comparisons put several Chinese open-weight systems ahead, and some of Mistral's strongest headline benchmark numbers are still based on vendor-published results from preliminary evaluations.

The most accurate verdict today is narrower and more interesting:

Large 4 appears to be the strongest open-weight model from Mistral so far, and the strongest non-Chinese open-weight model in the current comparisons, but the global open-model crown remains out of reach for now.

For developers, it is worth trying now if you care about long-context workloads, multilingual applications, multimodal reasoning, coding agents, or defensive cybersecurity tooling.

For companies considering production or self-hosted deployment, the smarter move is to evaluate the preview while waiting for three things to become clear: the final weights, the final license, and the final pricing.

Those three details may matter more than another few points on a benchmark leaderboard.


Sources

Facts and benchmark snapshots in this article are current as of October 6, 2026. Mistral-reported results are identified as such; third-party results are attributed to the relevant evaluator.

Comments (1)

Join the discussion by logging into your account.

Igor Ganapolsky

The 4.7 percent active figure is the right first cut, but it understates cost on an agent loop for two reasons that stack. Every expert still has to stay resident unless you offload, and offload latency sits on the critical path of each decoded token, so a tool-calling loop pays that stall once per step rather than once per request. Separately, about 200 million output tokens against an 81 million median is roughly 2.5 times the tokens on a comparable task. At the live 2.09 dollars per million output tokens, that verbosity tax is larger than the gap between the documented rate and the crossed-out 4.18 dollar rate. I would budget this preview from measured output length, and I would not treat today's numbers as the October checkpoint, because the reinforcement-learning run is still in flight.

Ankit Singh
Ankit Singh

Software Engineer

Passionate developer sharing knowledge about modern web technologies and best practices.

Subscribe to Ankit Singh's Newsletter

Direct email dispatches when new stories are published. Zero algorithms.