Mistral Large 4: Best Outside The US And China — And Still Not A Model To Run Your Agents On
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Mistral Large 4: Best Outside The US And China — And Still Not A Model To Run Your Agents On on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4 scored 38.4 on the Artificial Analysis Intelligence Index, a major improvement over its predecessor and the highest score among models from outside the United States and China in the supplied comparison. It still trails leading US and Chinese models, while its cost per benchmark task and reported verbosity raise questions about using it for agent workflows.

Mistral has released Large 4, a research-preview model that scored 38.4 on the Artificial Analysis Intelligence Index, according to the independent benchmark data cited in the source report. The result makes it the highest-scoring model from outside the United States and China in that comparison, but it remains below the leading US and Chinese systems—an important distinction for buyers considering it for demanding agent-based work.

Artificial Analysis Index version 4.3.2 puts Large 4 behind the listed US frontier models and several Chinese systems. Anthropic’s Claude Opus 5.5 scored 57.6, while China’s GLM-5.3 scored 44.8 and Kimi K3 scored 43.6. Large 4’s 38.4 places it ahead of some older or smaller systems in the table, but below seven Chinese open-weight models identified by the source report as its eventual open-weights comparison set.

The launch marks a sizable step up within Mistral’s own lineup. The source report says Mistral Large 3 scored 9 on the same index version and Medium 3.5 scored 14. The new model’s result therefore represents a substantial gain, although one benchmark cannot establish how it will perform across every task or deployment.

Mistral describes Large 4 as a one-trillion-parameter model with 49 billion active parameters, accepting text and images and producing text, with a 512,000-token context window. It is available through Mistral’s API as a research preview. The source says Mistral has promised to release its weights by the end of October; until that release, the model is proprietary and its licence has not been published. Listed pricing is $1.36 per million input tokens and $4.18 per million output tokens, with cached input priced at $0.14. A 50% discount is offered for the first two weeks, according to the supplied material.

At a glance
reportWhen: Released the day before the source repo…
The developmentMistral released Large 4 as a research preview, with benchmark data showing a sharp improvement over its predecessor but a continued gap to leading US and Chinese models.
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

The Agent-Work Cost Question

The benchmark matters because the Artificial Analysis Intelligence Index includes tasks intended to measure agentic work, including knowledge work, software workflows and coding. A model used in a multi-step process can carry an error forward into later actions, so performance on such tasks matters beyond how well it answers a single prompt. The source report argues that Large 4’s gap to higher-scoring systems, combined with its output volume, could make it a less attractive choice for long-running agent tasks.

Artificial Analysis data cited by the report shows Large 4 generated 200 million output tokens across the index tasks, compared with a median of 81 million for comparable models. The report also puts its cost at $1.13 per Intelligence Index task, against $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. Those two models scored 41.8 and 39.5, respectively. These are benchmark-task costs, not a guarantee of what a particular customer will pay: actual costs depend on workload, prompt length, caching and usage.

The practical question for organisations is not simply whether Large 4 is a strong European model. Its score suggests meaningful progress, but buyers weighing it against available alternatives will also need to test accuracy, latency, token use and cost on their own tasks. The source report author says hands-on testing found confident false claims; that is a personal observation, not a published Artificial Analysis result. It is a reason to evaluate reliability, not proof that the model will fail in every workflow.

Amazon

AI language model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A European Model in the Rankings

The headline claim that Mistral is home to the most intelligent model outside the US and China reflects a narrow geographic comparison. In the supplied index table, Large 4 leads models from elsewhere, while US systems occupy the top positions and multiple Chinese models score above it. The source report describes that outside-US-and-China race as having few direct contenders at this performance level. The benchmark supports the ranking within its listed models; it does not by itself show that Mistral leads every non-US or non-Chinese model on every capability.

Large 4’s score also arrives while its development remains active. According to the source material, Mistral says reinforcement learning is still underway and that scores may change. Its current availability as a preview means the published ranking should be read as a measurement of this version, rather than a final assessment of a finished model. The promised weights could broaden access to testing and deployment, but the terms of that release are not yet available in the supplied material.

“I saw something most of last month’s releases had stopped doing: confident hallucination.”

— ThorstenMeyerAI.com report author

Preview Results and Reliability

Several practical details remain unsettled. The source material does not provide the terms of the promised weights release, including the licence, and it does not say whether the end-of-October target has since been met. Mistral’s statement that reinforcement learning is ongoing also means the model’s benchmark score may change. The supplied report does not include a later, independently verified index result.

The reported hallucination concern is based on the author’s hands-on observations, without a stated sample size, testing protocol or measured rate. It should not be treated as a general error-rate estimate. Nor does the index alone settle how Large 4 compares on individual customer workloads. The material also does not supply a complete set of cost and quality results for every rival or explain how benchmark-task costs translate to production deployments.

Weights and Independent Retesting

The next stated milestone is Mistral’s promised weight release by the end of October. If it proceeds, the release should clarify the model’s licence and allow a wider range of users to examine and run the weights, subject to those terms. Mistral’s ongoing reinforcement learning could also lead to updated performance figures.

For now, buyers can access the research preview through Mistral’s API and compare it against alternatives on representative tasks, with attention to accuracy, output volume and total cost. Further index measurements and transparent testing of hallucination and agent reliability would help establish whether Large 4’s strong improvement over earlier Mistral models translates into dependable performance in production.

Key Questions

What is Mistral Large 4?

Mistral Large 4 is Mistral’s research-preview model. The source material describes it as a one-trillion-parameter, natively multimodal system with 49 billion active parameters and a 512,000-token context window.

How did it score against other models?

It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2. That is the highest listed score from outside the US and China in the supplied comparison, but several US and Chinese models scored higher.

Can users download its weights now?

Not according to the supplied report. Large 4 is currently offered as a proprietary API research preview, and Mistral has promised weights by the end of October. The licence for that release was unpublished in the source material.

Is Large 4 suitable for agent workflows?

The benchmark includes agent-oriented tasks, but the score alone does not determine whether the model is suitable for a particular workflow. The report raises concerns about its relative score, token use and observed hallucinations; organisations should test it on their own tasks before relying on it.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

ESPYS 2024: Unveiling Sports Excellence Celebration

Witness the pinnacle of sports excellence at ESPYS 2024, where champions shine and diversity thrives, leaving viewers captivated by unforgettable moments.

Steal This: The Signature Technique Behind “The Cenotaph At Night | A Two-Minute Vigil”

Exploring the signature design and technical approach of the virtual memorial ‘The Cenotaph at Night | A Two-Minute,’ emphasizing restraint and silence.

Some Claude Users Are Mad That Anthropic’s New Watermarks Will Catch Them Using It At Their Jobs, Classes – TechCrunch

Anthropic’s new watermarks in Claude models aim to ensure transparency but face criticism from users concerned about detection and privacy.

Janel Parrish Unveils Rich Cultural Heritage

Subtly weaving her diverse background into her performances, Janel Parrish showcases the beauty of cultural authenticity in Hollywood.