Mistral Large 4 Explained: Inside Europe's 1 Trillion Parameter Open-Weight Model (2026)
Written by
Aerin Kim

Mistral Large 4 is Europe's 1 trillion parameter open-weight MoE, trained on about 3,800 GPUs with a cyber-defense focus. See the benchmarks, pricing, release timeline, and runnable code.
If you run security operations, a legal team, or a regulated engineering group in Europe, you have probably hit the same wall: the strongest models live behind US or Chinese APIs, and the open-weight ones that fit your data-residency rules were a generation behind. On October 6, 2026, Mistral AI tried to change that with Mistral Large 4, a 1 trillion parameter, natively multimodal mixture-of-experts model that is available today as a public preview API, with open weights promised for the end of October [1].
This post goes through what Mistral actually announced, what independent measurements from Artificial Analysis add, and where the two disagree. It also covers why the model was trained on only about 3,800 GPUs, why Mistral is leaning so hard on cyber defense, and why shipping an API preview weeks before the weights is a release pattern worth understanding. You will also get runnable Python and bash snippets for calling the model, sending images, estimating cost per task, and sizing hardware for self-hosting.
TL;DR: What Mistral Large 4 Is in 2026
- Size: 1 trillion total parameters, with roughly 49 billion active per token (52 billion if you count embeddings and output layers) [1] [4].
- Type: an open-weight hybrid instruct and reasoning mixture-of-experts model that accepts text and images [1].
- Training: from scratch on 3,800 NVIDIA Grace Blackwell GPUs inside Mistral's own European datacenters [1].
- Languages: 160+ languages, including every official EU language [1].
- Price: $1.36 per million input tokens and $4.18 per million output tokens [1].
- Status: public preview API now on Mistral Studio, weights by the end of October 2026, with a Hugging Face placeholder listing an October 31 ETA [1] [4].
- Headline claims: 93% on Cybench, 82% on a reproduce-and-patch vulnerability test, 61.7% on DeepSWE v1.1, and 93.3% resistance on the Lakera B3 prompt-injection benchmark, all as reported by Mistral [1].
- Independent view: Artificial Analysis gives it an Intelligence Index of 38, a Cyber Index of 50, and 82% on CyberGym-E2E-AA, but flags a cost per task over four times that of similar-intelligence open models [2].
If you build content or products on top of language models, the same pipeline thinking applies to media. Miraflow's AI image generator, cinematic video tool and home workspace are where many teams turn a research summary like this one into diagrams and short explainers.
Why Mistral Large 4 Matters: The European Sovereign Open-Weight Frontier
The open-weight frontier in 2026 has been crowded, but almost entirely with Chinese and a few American labs. If you read our breakdowns of Kimi K3, Tencent Hy4 Preview, Qwen3.8-Max and Reflection AI's Beam, the pattern is clear: very large sparse models, mostly released by non-European labs, with deployment terms that some European buyers cannot accept.
Mistral Large 4 is positioned as the European answer. TechCrunch reported that Mistral hopes it will be the best open-weight model, especially outside China, and quoted Emmanuel Macron calling it "a third way in AI" [3]. Mistral's own announcement is specific about the sovereignty angle: the preview runs on the same European infrastructure used for training, a European deployment option is operated end to end under European law, and organizations can self-deploy on private cloud or on-premise hardware [1].
That matters for a practical reason rather than a political one. Security operations teams, banks, public agencies and defense suppliers often cannot send logs, malware samples or contracts to a US-hosted API. The EU's regulatory framework for AI also places obligations on general-purpose model providers, which pushes procurement teams toward vendors with a clear European legal footprint [19]. A strong model that can run inside your own perimeter removes a whole category of review.
Where this sits against Mistral's own history
Mistral Large 3 shipped in December 2025 as a sparse mixture-of-experts model with 41 billion active and 675 billion total parameters, released under Apache 2.0 [7]. Large 4 grows total parameters by roughly 48% (675B to 1T) and active parameters from about 41B to about 49B to 52B depending on how you count [1] [4] [7]. The documentation overview lists Mistral Large 4 as version v26.10 with an "Open" license label, while Large 3 (v25.12) is listed as Apache 2.0 [6]. The exact license terms for Large 4 had not been published at the time of writing, so check the Hugging Face model card on release day before you build a commercial product around the weights.
Large 4 also arrives after Mistral raised a €3 billion Series D, which Mistral describes as the largest equity round ever raised by a European technology company, and calls ML4 the first milestone that funding paid for [1]. TechCrunch adds that Samsung led the Series D at a €21 billion valuation and that ASML led the earlier Series C, which helps explain why chip design shows up in the use-case list [3].

Inside the Architecture: 1 Trillion Parameters, About 5% Active
A mixture-of-experts layer replaces one big feed-forward block with many smaller expert blocks and a router that sends each token to only a few of them. The idea goes back to Shazeer et al.'s sparsely-gated MoE layer [8], was scaled to trillion-parameter models in Google's Switch Transformer [9], and went mainstream for open weights with Mixtral [10] and DeepSeek-V3 [11]. Hugging Face's explainer is a good primer if the router concept is new to you [18].
For Mistral Large 4 the numbers work out as follows:
- 49B active out of 1,000B total is 4.9% of parameters used per token.
- 52B active including embeddings and output layers is 5.2%.
- In plain terms, if you laid out 1,000 tokens that each represent a billion parameters, about 52 of them would be lit up for any given word.
The reason to accept this design is cost per token. Compute per generated token scales with active parameters, so a 1T-parameter MoE with 49B active does roughly the arithmetic of a 49B dense model for each token. What it does not save is memory: every expert must be resident somewhere, which is why the self-hosting section below starts from 1 TB of weights at FP8.
What Mistral has and has not disclosed
Mistral says the weights release will come with more detail on architecture, benchmarks and post-training [1]. That means several things you would normally want are still unknown as of October 10, 2026:
- The number of experts, how many are routed per token, and whether there are shared experts.
- The attention design and how it handles long sequences.
- The tokenizer size and the vision encoder design.
- The exact context window. Artificial Analysis lists 512K tokens [2], some coverage mentions 1M, and Mistral's announcement and docs overview do not state a figure [1] [6]. Treat any number as provisional until the model card is published.
This is a good moment to say plainly what is verified and what is not. Total and active parameter counts, the GPU count, the price, and the release timing come from Mistral and are corroborated by Hugging Face, TechCrunch and Artificial Analysis. Everything about the internal layer structure is open.
The hybrid instruct and reasoning design
Earlier Mistral generations split chat-style and reasoning-style behavior across different checkpoints. Mistral describes Large 4 as a single hybrid model that unifies instruction following, reasoning and agentic work [1]. For application developers that removes a routing decision. If you have built a router between a fast model and a thinking model, our walkthrough of model routing with NVIDIA NeMo Switchyard and Runway's router shows what that layer looks like and when a single hybrid model lets you delete it.
Multimodal input and the 100-image change
Large 4 accepts images as input and produces text. Artificial Analysis notes that the API now allows 100 images per request, up from 8 [2]. That is a bigger change than it sounds for document work: a 60-page scanned contract or a multi-sheet engineering drawing set fits in one call instead of being chunked. Mistral's own visual grounding result on Dense 200 is 42% against 41% for GPT-6-Astra [1], and TheNextWeb reports 73% versus 68% on DIOR-RSVG, an earth observation grounding set [5]. Those are vendor-reported and preliminary, but they show where the training effort went.
Training on About 3,800 GPUs: The Efficiency Story
Mistral says Large 4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacenters [1]. TechCrunch rounds that to about 4,000 and quotes Pierre Stock, Mistral's VP of Science, saying this is two to three times fewer than Chinese competitors and far fewer than closed-source rivals [3]. TheNextWeb adds that the run took roughly two months and drew about 10 megawatts [5].
Here is a back-of-envelope reading of those figures, clearly labeled as my arithmetic rather than a Mistral disclosure:
- If 3,800 GPUs is two to three times fewer, the comparison fleets would be roughly 7,600 to 11,400 GPUs.
- Two months is roughly 1,460 hours, so 3,800 GPUs for that long is on the order of 5.5 million GPU-hours.
- 10 megawatts across 3,800 GPUs is about 2.6 kilowatts per GPU once you include CPUs, networking, cooling overhead and storage.
None of those numbers is a precise accounting, and Mistral has not published a token count for pretraining. They are still useful for calibrating how a lab with a smaller fleet can field a 1T-parameter model. The main reasons, which Mistral and independent MoE research both point to, are sparse compute per token, the FP8 and lower-precision training formats that Blackwell-class hardware supports [13], and a post-training stack that spends compute where it moves benchmarks.

Why a smaller GPU count does not mean a weaker model
Pretraining cost for a MoE scales with active parameters times tokens, not total parameters times tokens. DeepSeek-V3 made this argument publicly in late 2024 with a 671B-parameter, 37B-active model trained at modest cost by the standards of the time [11]. Mistral is making the same argument at a larger size and on newer hardware. The honest caveat is that fewer GPUs usually means either fewer pretraining tokens or a longer calendar time, and Mistral has said the reinforcement learning run is still in flight and expects large improvements as it progresses [1]. That is also why the Hugging Face release might score higher than the preview you can call today.
The RL stage in numbers
The most detailed engineering disclosure in the announcement is about reinforcement learning, so it deserves a close read [1]:
- Compute: RL runs on about 3,000 GPUs.
- Throughput: one run produces about 33 billion tokens per day, of which about 16 billion are trainable completion tokens.
- Rollouts: an autoscaling fleet generates tens of thousands of rollouts in parallel while training runs asynchronously, with rollout budgets reaching millions of tokens across multiple context compactions.
- Verification: reward models, unit tests, LLM judges and static checks, combined behind one shared interface that covers single-turn chat, scientific problem solving, safety alignment, factuality and long-horizon tool use.
Divide 33 billion tokens per day by 3,000 GPUs and you get roughly 11 million tokens per GPU per day, or about 127 tokens per second per GPU, most of it spent generating rollouts rather than updating weights. About 48% of generated tokens end up as trainable completions, which tells you how much of the RL budget goes to context, tool outputs and discarded attempts. If you want the background on why labs build thousands of verifiable tasks for this stage, see our explainer on environment scaling in RL training.
Mistral also says the training, customization and RL environment used for Large 4 is the same one it offers customers through Mistral Forge [1]. For an enterprise buyer, that is the most commercially interesting sentence in the announcement: the recipe that produced the flagship is the recipe you can point at your own tasks.
The Preview-Before-Weights Release Pattern
Large 4 is the clearest recent example of a release pattern that is spreading: ship a hosted preview first, publish weights later. The calendar looks like this:
| Date | Event | Source |
|---|---|---|
| Oct 6, 2026 | Public preview API on Mistral Studio, announcement | Mistral |
| Oct 6 to about Oct 20 | 50% launch discount per Artificial Analysis | Artificial Analysis |
| Oct 27, 2026 | Weights date reported by TheNextWeb | TheNextWeb |
| Oct 31, 2026 | Current ETA on the Hugging Face placeholder | Hugging Face |
| End of Oct 2026 | Weights per Mistral's own announcement | Mistral |
The Oct 27 versus Oct 31 gap is a real discrepancy between sources. Mistral itself says only "by the end of the month" [1], TechCrunch says about three weeks after the October 6 announcement [3], and the Hugging Face page shows a countdown to October 31 [4]. Plan for the later date.

Why ship the API first
Mistral gives the safety reason directly. TechCrunch quotes Stock saying weights follow once safety testing is complete, and that Mistral will work with trusted partners and governments so the open weights can be used for defense rather than attacks [3]. Mistral's announcement says it is red-teaming with cybersecurity leaders, vetted partners and state authorities, who get a version with reduced moderation and expanded cyber capabilities [1]. TheNextWeb describes that tier as going to developers, cybersecurity firms and government agencies [5].
There are also engineering reasons. An API preview gives Mistral live traffic to catch regressions while the RL run continues, and it lets third parties like Artificial Analysis publish independent numbers before weights exist, which is how this launch got a Cyber Index score on day one [2]. For you as a builder, the preview is a cheap way to test whether the model fits your workload before you commit to GPUs.
What a preview means for your integration plan
- Pin and record the version. The documentation lists Large 4 as v26.10 [6]. Log the exact model string in every evaluation run, because the RL stage is ongoing and behavior may shift.
- Expect benchmark drift. TheNextWeb notes that benchmark results are preliminary and Mistral expects them to change before the weights are released [5].
- Do not hardcode the weights schedule. Build with the API now and make self-hosting a second milestone.
- Check the license on the day. The docs label it "Open" without naming a license [6].
Cyber Defense: The Focus That Sets Large 4 Apart
Most launch posts in 2026 treat cybersecurity as one benchmark row. Mistral treats it as the headline. Its use-case list starts with vulnerability research, incident response, malware analysis and writing detection rules [1]. This lines up with a broader 2026 trend we covered in Gemini 4 Argon, which Google restricted to cyber defenders, and with the way GPT-6 Astra reached a critical cyber threshold.
The cyber numbers
| Benchmark | Mistral Large 4 | Source |
|---|---|---|
| Reproduce and patch a real vulnerability | 82% | Mistral |
| CyberGym-E2E-AA | 82% (MiMo-V2.6-Pro 79%, GPT-6 Luna max 78%) | Artificial Analysis |
| Cybench (40 challenges) | 93% | Mistral |
| Artificial Analysis Cyber Index | 50 (MiMo-V2.6-Pro 56, GLM-5.3-Flash 50) | Artificial Analysis |
| Lakera B3 prompt-injection resistance | 93.3% | Mistral |
The 82% figure appears in both sources and probably refers to the same underlying test, which is built on CyberGym, a benchmark of real-world vulnerabilities where an agent must reproduce a bug from a description [17]. Cybench is an academic framework of professional-level capture-the-flag tasks [12]. Mistral says it solves 93% of the 40 challenges, one of the highest scores reported for an open-weight model [1].
Now the nuance. Artificial Analysis says Large 4 would rank among the top three open-weight models on its Cyber Index once weights are released, but its index score of 50 is level with GLM-5.3-Flash and behind MiMo-V2.6-Pro at 56 [2]. Mistral's phrasing is that it is among the top five models globally and leads open-weight models developed outside China by a wide margin [1]. Both can be true. The strongest honest summary is that Large 4 is the best non-Chinese open-weight cyber model so far, and a very strong performer on end-to-end vulnerability reproduction, but not the top of the full Cyber Index.

The refusal gap
One claim deserves special attention. Mistral says several closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on the vulnerability reproduction test because they refuse the task [1]. If that holds up, it exposes a real operational problem: a defender asking a safety-tuned model to reproduce a bug in their own code is indistinguishable, at the prompt level, from an attacker asking the same thing. Mistral's answer is a tiered access model, with a reduced-moderation version for vetted partners and a standard version for everyone else [1] [5].
I would read this carefully. A near-zero score from a refusing model measures policy, not capability, and the comparison is Mistral's own. It is still a useful signal that cyber teams hitting refusals on hosted APIs now have a credible alternative they can run under their own controls. For a view of what happens when agents get that capability in the wild, see our write-up of the PaperCut attack that breached 395 organizations.
Prompt injection and robustness
Security teams care about defending the model as much as using it. Large 4 resists 93.3% of attacks on Lakera's public B3 benchmark, and Mistral reports 1.691 out of 2 on KORA, which it says is the highest measured among open-source models [1]. Indirect prompt injection, where a malicious instruction hides in a web page, email or tool output, is ranked first in the OWASP Top 10 for LLM applications [14]. Our analysis of ToolHazard and indirect prompt injection shows why a single benchmark number never replaces testing with your own tools.
Worked example: a detection-rule workflow
Mistral says internal testing found the model useful for malware analysis, vulnerability prioritization and writing detection rules [1]. A realistic workflow with the API looks like this:
- Feed the model a sanitized indicator report and a handful of log lines.
- Ask for a Sigma or YARA rule plus a list of expected false positives.
- Run the rule against a test corpus and feed the failures back for a second pass.
- Keep a human reviewer on any rule that will block production traffic.
For samples you cannot send to a hosted API, this is where the self-deploy option on private cloud or on-premise hardware earns its keep [1].
Benchmarks Beyond Cyber: Coding, Agents, Documents
Cyber is the focus, but buyers will evaluate the model on general work. The table below collects Mistral's reported numbers with the independent ones next to them.
| Benchmark | Mistral Large 4 | Comparison / note | Source |
|---|---|---|---|
| DeepSWE v1.1 | 61.7% | GLM-5.3 61%, DeepSeek-V4-Pro 57% (TheNextWeb) | Mistral |
| SWE-Atlas-QnA | 59.4% | Vendor-reported | Mistral |
| Terminal-Bench 4 | 28.3% | Vendor-reported | Mistral |
| Coding Agent Index | 49.8% | Ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max | Mistral |
| Surge AI blind coding eval (1 to 5) | 3.74 (2nd of 5) | Claude Opus 5 4.22, GLM-5.3 3.60, Kimi K3 3.59, GLM-5.2 3.40 | Mistral |
| AutomationBench (657 workflows) | 59.9% | Ahead of Kimi K3, MiMo-V2.6-Pro, DeepSeek V4 Pro | Mistral |
| AA-Briefcase | 1,393 Elo | Ahead of DeepSeek V4 Pro | Mistral |
| Dense 200 visual grounding | 42% | GPT-6-Astra 41% | Mistral |
| GDP.pdf | 19% | Kimi K3 22%, MiMo-V2.6-Pro 19% | Artificial Analysis |
| Intelligence Index | 38 | GPT-6 Luna (max) 38, DeepSeek V4.1 Flash (max) 39, Claude Haiku 5.5 43 | Artificial Analysis |
Coding
DeepSWE v1.1 at 61.7%, SWE-Atlas-QnA at 59.4% and Terminal-Bench 4 at 28.3% are the vendor-reported coding results [1]. TheNextWeb reports DeepSWE as 62%, narrowly ahead of GLM-5.3 at 61% and DeepSeek-V4-Pro at 57% [5]. Mistral also reports a Coding Agent Index of 49.8%, ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max [1]. For context on how Terminal-Bench scores are produced and why they vary by harness, read our Terminal-Bench 4.0 explainer, and for the competing coding releases see DeepSeek V4-Pro 0813 and GLM-5.3.
The most interesting coding data point is human rather than automated. In a blind Surge AI coding evaluation on a 1 to 5 scale, the Large 4 preview scored 3.74, second of five models, behind Claude Opus 5 at 4.22 and ahead of GLM-5.3 at 3.60, Kimi K3 at 3.59 and GLM-5.2 at 3.40 [1]. A gap of 0.48 to the leader is real, and it is a reminder that open-weight models are close behind the best closed ones, not level with them.

Agentic and knowledge work
AutomationBench, which covers 657 business workflows, comes in at 59.9%, ahead of Kimi K3, MiMo-V2.6-Pro and DeepSeek V4 Pro, and AA-Briefcase gives 1,393 Elo [1]. Mistral says Large 4 exceeds GPT-6-Astra on vals.ai's Finance Agent v2 and outperforms all open-source models on Harvey's Legal Agent Benchmark, though it does not print the figures in the text [1]. TheNextWeb fills in some of the gaps: 67% on FinWorkBench, level with DeepSeek-V4-Pro, and 15% on Harvey Legal Agent versus 13% for Kimi K3 and 5% for GPT-6 Astra, with the reminder that legal scores are low for every model [5].
Documents and images
Artificial Analysis's GDP.pdf result is 19%, an 18-point improvement over Mistral Large 3, level with MiMo-V2.6-Pro and behind Kimi K3 at 22% [2]. So the multimodal jump is large relative to Mistral's own past, and still not the best among open models. That is consistent with the way Mistral talks about the model: strong on specific verticals, not best everywhere.
What Independent Measurements Say
Artificial Analysis evaluated the preview and gave it an Intelligence Index of 38, comparable to GPT-6 Luna at max effort (38) and DeepSeek V4.1 Flash at max effort (39), and described it as the most intelligent model from outside the US and China [2]. For comparison, Claude Haiku 5.5 scores 43 on the same index [2].
The less flattering finding is cost. Artificial Analysis puts the cost to run its Intelligence Index at $1.13 per task at standard pricing, which it says is over four times the cost of similar-intelligence open-weight models, versus $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash [2]. The 50% launch discount for the first two weeks ($0.68 input and $2.09 output) lowers that to $0.57 per task [2].
| Pricing view | Input ($/1M tokens) | Output ($/1M tokens) | Cost per Intelligence Index task |
|---|---|---|---|
| Standard | $1.36 | $4.18 | $1.13 |
| Launch discount (first two weeks) | $0.68 | $2.09 | $0.57 |
| Cached input (standard) | $0.14 | n/a | n/a |
| GLM-5.3-Flash (for comparison) | not covered here | not covered here | $0.25 |
| DeepSeek V4.1 Flash, max (for comparison) | not covered here | not covered here | $0.27 |
Reading the cost result fairly
A higher cost per task has two possible causes: a higher token price or more tokens consumed per task. Large 4's list price of $1.36 and $4.18 is not extreme, so the cost gap probably includes long reasoning traces, although Artificial Analysis did not report token usage in the article I read [2]. If you are comparing models for a production workload, measure tokens per completed task on your own data rather than trusting either list price or an index average.

The economics also change when weights land. Self-hosting turns a per-token price into a fixed cluster cost, and the right choice depends on utilization, as covered in the hardware section below. For a similar API-versus-weights comparison on another large MoE, our Kimi K3 explainer walks through the break-even logic, and the DeepSeek-V4.1-Flash explainer shows how a much cheaper model changes the calculation.
Calling Mistral Large 4 From Python
This section is the practical core. The snippets use the official mistralai Python SDK and read the key from an environment variable so no credential ever appears in code [16]. One honest warning applies to all of them: Mistral's announcement does not print the API model string for the preview, so I use the long-standing mistral-large-latest alias and you should confirm in the model list at docs.mistral.ai which alias currently points to Large 4 [1] [6]. The docs page slug suggests mistral-large-4-0, but that is not a confirmed model ID.
Install the SDK and set your key in your shell first:
bash/code # Install the official Mistral SDK and set your key for this shell session. # Never paste a real key into source files or commit it to git. pip install --upgrade mistralai export MISTRAL_API_KEY="paste-your-key-here-in-your-terminal-only" echo "Key set: ${MISTRAL_API_KEY:+yes}"
Example 1: a basic chat completion for security triage
python/code import os from mistralai import Mistral # The key comes from the environment, never from source code. client = Mistral(api_key=os.environ["MISTRAL_API_KEY"]) # Confirm at docs.mistral.ai which alias points to Mistral Large 4. # "mistral-large-latest" is the long-standing alias; the exact ID may differ. MODEL = "mistral-large-latest" response = client.chat.complete( model=MODEL, temperature=0.2, messages=[ { "role": "system", "content": "You are a SOC analyst assistant. Be precise, list assumptions, and flag anything that needs human review.", }, { "role": "user", "content": ( "Triage this log line and suggest one detection idea:\n" "powershell.exe -nop -w hidden -enc SQBFAFgA... launched by winword.exe" ), }, ], ) print(response.model) # log the exact model string for reproducibility print(response.choices[0].message.content) print(response.usage) # prompt_tokens, completion_tokens, total_tokens
Setting temperature low keeps triage answers consistent between runs. Since the preview model can change while RL continues, log response.model and a hash of your prompt with every call so you can explain a changed result later.
Example 2: sending an image
The 100-images-per-request limit that Artificial Analysis reports makes multi-page documents practical [2]. The next snippet sends one local screenshot as a base64 data URL. If you need many pages, append more image items to the same content list.
python/code import base64 import os from mistralai import Mistral client = Mistral(api_key=os.environ["MISTRAL_API_KEY"]) MODEL = "mistral-large-latest" # verify the alias for Large 4 in the docs def to_data_url(path: str) -> str: with open(path, "rb") as f: encoded = base64.b64encode(f.read()).decode("utf-8") return f"data:image/png;base64,{encoded}" content = [ {"type": "text", "text": "Read this architecture diagram and list every network trust boundary it shows."}, {"type": "image_url", "image_url": to_data_url("diagram.png")}, # Add more image_url items here for multi-page documents. # Artificial Analysis reports a limit of 100 images per request. ] response = client.chat.complete( model=MODEL, messages=[{"role": "user", "content": content}], ) print(response.choices[0].message.content)
Example 3: a cost-per-task calculator
List prices are easy to misread when a task involves thousands of input tokens and a few hundred output tokens. This script uses the real preview prices from the announcement and Artificial Analysis, including the two-week launch discount and the cached-input rate [1] [2].
python/code # Cost per task for Mistral Large 4 preview pricing. # Prices per 1M tokens: $1.36 input, $4.18 output (Mistral announcement). # Cached input $0.14 and a 50% launch discount for two weeks (Artificial Analysis). INPUT_PER_M = 1.36 OUTPUT_PER_M = 4.18 CACHED_PER_M = 0.14 def task_cost(input_tokens, output_tokens, cached_tokens=0, discount=0.0): fresh = input_tokens - cached_tokens cost = ( fresh * INPUT_PER_M + cached_tokens * CACHED_PER_M + output_tokens * OUTPUT_PER_M ) / 1_000_000 return cost * (1 - discount) input_tokens, output_tokens = 20_000, 3_000 tasks_per_day = 10_000 for label, kwargs in [ ("standard", {}), ("launch discount (50%)", {"discount": 0.5}), ("standard, 15k tokens cached", {"cached_tokens": 15_000}), ]: c = task_cost(input_tokens, output_tokens, **kwargs) print(f"{label:32s} ${c:.4f} per task, ${c * tasks_per_day:,.0f} per day")
With the sample numbers in the script, a 20,000-token input and 3,000-token output costs about $0.04 at standard pricing and about $0.02 during the launch discount. At 10,000 such tasks per day that is roughly $397 versus $199 per day, which is the kind of calculation that decides whether a pilot becomes a product.
Self-Hosting a 1 Trillion Parameter Model
Weights are not out yet, so everything here is planning rather than a tested deployment. The first constraint is memory: every expert must be loaded even though only about 5% fire per token.

| Precision | Bytes per parameter | Weights only (1T params) | Typical note |
|---|---|---|---|
| BF16 | 2 | about 2 TB | Reference quality, rarely worth it for serving |
| FP8 | 1 | about 1 TB | Common serving target on Hopper and Blackwell |
| FP4 | 0.5 | about 0.5 TB | Needs Blackwell-class support and quality checks |
On top of weights you need KV cache, activations and headroom. Because the context window is unconfirmed, I would not size KV cache until the model card appears. The snippet below computes weight memory and a minimum GPU count for a few accelerator sizes, with a 20% overhead assumption you should replace with measured numbers.
python/code import math # Back-of-envelope weight memory for a 1T-parameter MoE. # All experts must be resident even though only ~49B parameters are active per token. TOTAL_PARAMS = 1_000_000_000_000 ACTIVE_PARAMS = 49_000_000_000 OVERHEAD = 1.20 # assumed 20% for KV cache, activations, buffers; measure your own bytes_per_param = {"BF16": 2.0, "FP8": 1.0, "FP4": 0.5} gpu_memory_gb = {"80 GB class": 80, "141 GB class": 141, "192 GB class": 192} print(f"Active fraction: {ACTIVE_PARAMS / TOTAL_PARAMS:.1%}") for prec, b in bytes_per_param.items(): weights_gb = TOTAL_PARAMS * b / 1e9 need_gb = weights_gb * OVERHEAD print(f"\n{prec}: weights {weights_gb:,.0f} GB, with overhead {need_gb:,.0f} GB") for name, gb in gpu_memory_gb.items(): print(f" min GPUs at {name}: {math.ceil(need_gb / gb)}")
Planning checklist for release day
- Confirm the license, context length, expert count and recommended inference engine on the Hugging Face model card [4].
- Check that your serving stack supports the architecture. vLLM is the usual first target for new open-weight MoE models [15].
- Plan expert parallelism across GPUs, because a single 8-GPU node cannot hold 1 TB at FP8 unless each GPU has well over 128 GB.
- Budget network bandwidth between nodes, since token routing across experts adds all-to-all traffic.
- Download only after you have disk space. When the repository goes live, the command below is what you would use; the repo name comes from the placeholder page and may change.
bash/code # Run only after the weights are published and you have checked the license and disk space. # The repo name comes from the Hugging Face placeholder page and may change. pip install --upgrade huggingface_hub hf auth login # paste your token interactively, never inline df -h . # expect on the order of 0.5 to 2 TB depending on the published precision hf download mistralai/Mistral-Large-4-1T-A52B --local-dir ./mistral-large-4
When self-hosting actually makes sense
For most teams the hosted API is the better first step. Self-hosting wins when you have sustained, predictable volume, strict data residency, or a need to run the reduced-moderation cyber tier inside a controlled network. Mistral's pitch to security operations teams is exactly that last case [1]. If your load is bursty, rent API capacity, and revisit when you can keep GPUs busy.
How Large 4 Compares With Other Open-Weight Frontier Models
No single table captures all differences, but a compact comparison helps you decide where to look next.
- Versus Kimi K3: Kimi K3 leads Large 4 on GDP.pdf (22% versus 19%) [2], while Large 4 leads Kimi K3 on the Surge blind coding rating (3.74 versus 3.59) and on AutomationBench [1]. Kimi K3 is much larger at 2.8 trillion parameters, as covered in our Kimi K3 post.
- Versus Reflection AI's Beam and Tencent Hy4: both are open-weight MoE models in the 500B to 770B range, covered in the Beam explainer and the Hy4 explainer. Large 4 is larger and has a stronger cyber story, while those posts show different strengths in agentic training.
- Versus GLM-5.3: GLM-5.3-Flash matches Large 4 on the Cyber Index at 50 and costs $0.25 per task versus $1.13 [2]. If cost per cyber task is your metric, that comparison is uncomfortable for Mistral until the discount or weights change the math.
- Versus closed models: Claude Opus 5 leads the blind human coding rating, and the Large 4 preview trails it by 0.48 points [1]. Closed models are also gated differently, as the cyber refusal discussion above shows.
Common Mistakes When Evaluating a Preview Model
- Treating vendor benchmarks as final. The RL run is still in flight and TheNextWeb says results are preliminary [1] [5].
- Mixing parameter counts. 49B and 52B active both appear, depending on whether embeddings and output layers are counted [4].
- Assuming a context window. Sources disagree, and Mistral does not state one [1] [2].
- Comparing price per token instead of cost per task. Reasoning-heavy models can cost more per task even at similar list prices [2].
- Ignoring the refusal comparison's source. The near-zero scores for closed models come from Mistral's announcement, so reproduce on your own tasks [1].
- Hardcoding model IDs and dates. Check the docs and the model card.
- Assuming self-hosting is cheaper. A 1 TB model needs a multi-node cluster before the first token is served.
Turning Research Into Visuals With Miraflow
Technical releases like this are hard to explain in words alone, which is why the images in this post show the actual numbers on labeled props. If you write explainers, newsletters or internal briefings, you can do the same with Miraflow. The AI image generator handles tabletop-style infographics and labeled diagrams, the YouTube thumbnail maker helps you package a model-launch video, and Text2Shorts turns a script into a vertical summary for Shorts, Reels and TikTok. If you record a long walkthrough of a launch, AI Clipping finds the best moments and cuts captioned clips, and AI music can add a background bed.
Here is an image prompt you can adapt for a similar sparse-activation visual, written for a Nano Banana style photoreal model:
Photoreal tabletop still life on a dark walnut desk: a rectangular wooden tray holding exactly 1,000 small brass tokens arranged in a 40 by 25 grid, with 52 of the tokens polished bright copper and the rest dull brass, a small kraft paper tag in the corner reading "52 of 1,000 active". Soft window light from the left, camera 45 degrees above at a close distance with shallow depth of field on the copper tokens, visible metal texture and wood grain. Sharp, accurate details, no warped or blurry tokens, no garbled text, no watermark or logo.
And here is a video prompt written for a Wan-style text-to-video model, useful for a short explainer clip:
A slow 6 second dolly shot across a walnut desk lit by warm window light from the left. A miniature server rack with a small paper tag reading "3,800 GPUs" sits beside a wall calendar with the dates "Oct 6" and "Oct 31" circled in red pen. A hand turns a page of the calendar from the 6th to the 31st while small brass tokens on the desk light up copper one by one. Photoreal, shallow depth of field, steady camera, natural hands, no text distortion, no logos, no watermark.
You can generate clips like this with the cinematic video tool, and plan a full pipeline from idea to script to visual to video from the Miraflow home page.
What to Watch Before and After the Weights Drop
- Model card details. Expert count, context window, license and vision encoder are all still open [1] [4].
- Independent cyber results. Artificial Analysis says it expects to re-rank the model on release [2].
- RL progress. Mistral expects large improvements as the run proceeds, and says ML4 will be the base for specialized future models [1].
- Compute growth. TheNextWeb reports additional compute coming online through the first half of 2027, so the 3,800-GPU headline is likely a floor rather than a ceiling for Mistral's next run [5].
- Third-party hosting. Open weights usually lead to inference providers listing the model within days, which changes the price comparison.
For related coverage of open-weight launches that followed the same preview-first idea, see Atria Dawn Preview and Thinking Machines' Inkling, and for a hardware-side perspective on the same hybrid-architecture trend, NVIDIA's Nemotron 3 Ultra.
Frequently Asked Questions
Is Mistral Large 4 open source?
Mistral describes it as open-weight and says weights will be released by the end of October 2026 [1]. The documentation lists the license as "Open" without naming one, while Mistral Large 3 is Apache 2.0 [6]. Read the license on the release page before commercial use.
How many parameters does Mistral Large 4 have?
1 trillion total. About 49 billion are active per token, or 52 billion including embeddings and output layers [4].
How many GPUs was it trained on?
Mistral says 3,800 NVIDIA Grace Blackwell GPUs in its European datacenters [1]. TechCrunch rounds this to about 4,000 and quotes Mistral saying that is two to three times fewer than Chinese competitors [3].
When can I download the weights?
Mistral says by the end of the month [1]. Hugging Face lists an ETA of October 31, 2026 [4], and TheNextWeb reports October 27 [5].
How much does the API cost?
$1.36 per million input tokens and $4.18 per million output tokens [1]. Artificial Analysis lists $0.14 per million cached input tokens and a 50% launch discount for the first two weeks, which brings it to $0.68 and $2.09 [2].
What is the context window?
It is unconfirmed. Artificial Analysis lists 512K tokens [2], other coverage mentions 1M, and Mistral's announcement does not state a figure [1]. Check the docs.
Can it process images?
Yes, it accepts image input and outputs text, and the API reportedly accepts up to 100 images per request [2].
Is it good at cybersecurity?
Mistral reports 93% on Cybench and 82% on a vulnerability reproduce-and-patch test [1], and Artificial Analysis measured 82% on CyberGym-E2E-AA and a Cyber Index of 50, behind MiMo-V2.6-Pro at 56 [2]. It is very strong on vulnerability work and not the overall leader.
Can I run it locally?
Not until the weights are released, and even then you need roughly 1 TB of memory for weights at FP8, which means a multi-GPU cluster rather than a workstation.
Is Mistral Large 4 better than Kimi K3?
It depends on the task. Large 4 scores higher on the Surge blind coding evaluation (3.74 versus 3.59), while Kimi K3 scores higher on GDP.pdf (22% versus 19%) [1] [2].
Conclusion
Mistral Large 4 is the first European model to enter the open-weight trillion-parameter tier, and the interesting part is how it got there: about 3,800 GPUs, a sparse design that activates around 5% of its parameters per token, and an RL stage that is still running. The cyber results are strong enough to matter for defenders, especially those who hit refusals on closed APIs or cannot send data outside their own network. The preview-then-weights release lets you test it now and plan hosting later.
The caveats are equally real. Several key facts are still undisclosed or inconsistent across sources, including context length, license terms and the exact weights date. Artificial Analysis's cost-per-task figure says you should measure total task cost before assuming a mid-range token price means a cheap workload. Build against the API today, pin the version, run your own evaluations, and decide on self-hosting after the model card is out.
If you create explainers about releases like this, try the Miraflow AI image generator to turn the numbers into labeled visuals, or start from the Miraflow home page to plan the whole pipeline.
References
- [1] Mistral AI, "Mistral Large 4" (announcement, October 6, 2026). https://mistral.ai/news/mistral-large-4
- [2] Artificial Analysis, "Mistral Large 4". https://artificialanalysis.ai/articles/mistral-large-4-france-ai
- [3] TechCrunch, "Mistral's new 1T model aims to leapfrog closed and open rivals". https://techcrunch.com/2026/10/06/mistrals-new-1t-model-aims-to-leapfrog-closed-and-open-rivals/
- [4] Hugging Face, mistralai/Mistral-Large-4-1T-A52B (upcoming release page). https://huggingface.co/mistralai/Mistral-Large-4-1T-A52B
- [5] TheNextWeb, "Mistral releases Large 4, a 1 trillion parameter open-weight AI model". https://thenextweb.com/news/mistral-releases-large-4-a-1-trillion-parameter-open-weight-ai-model
- [6] Mistral AI documentation, Models overview. https://docs.mistral.ai/models/overview
- [7] Mistral AI, "Introducing Mistral 3" (Mistral Large 3, December 2025). https://mistral.ai/news/mistral-3
- [8] Shazeer et al., "Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer". https://arxiv.org/abs/1701.06538
- [9] Fedus et al., "Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity". https://arxiv.org/abs/2101.03961
- [10] Jiang et al., "Mixtral of Experts". https://arxiv.org/abs/2401.04088
- [11] DeepSeek-AI, "DeepSeek-V3 Technical Report". https://arxiv.org/abs/2412.19437
- [12] Zhang et al., "Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models". https://arxiv.org/abs/2408.08926
- [13] Micikevicius et al., "FP8 Formats for Deep Learning". https://arxiv.org/abs/2209.05433
- [14] OWASP Gen AI Security Project, "LLM01:2025 Prompt Injection". https://genai.owasp.org/llmrisk/llm01-prompt-injection/
- [15] vLLM documentation. https://docs.vllm.ai/
- [16] Mistral AI documentation. https://docs.mistral.ai/
- [17] Wang et al., "CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale". https://arxiv.org/abs/2506.02548
- [18] Hugging Face, "Mixture of Experts Explained". https://huggingface.co/blog/moe
- [19] European Commission, "AI Act". https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai


