🔍 Read the full analysis: Opus Builds, Sol Digs, Jev Decides — And Mistral Large 4 Doesn’t Make The Cut: My October 2026 AI Stack on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
ThorstenMeyerAI.com reports that Mistral Large 4 did not meet the author’s cost-and-capability threshold in a comparison using Artificial Analysis Intelligence Index v4.3.x scores and estimated task costs. The author says seven model configurations score higher and cost less per task, while noting Mistral’s speed, cyber benchmarks and expected European weights as potential advantages. These are one developer’s findings, not a universal assessment of model performance.
ThorstenMeyerAI.com reports that its author will not add Mistral Large 4 to a working AI model stack after comparing it with other systems on benchmark scores and estimated cost per task. In the author’s analysis, seven configurations score higher and cost less, a finding relevant to developers choosing models by both capability and operating expense. The comparison relies on a general benchmark and the author’s cost estimates, not tests of every real-world workload.
The report places Mistral Large 4 at 38 points on the Artificial Analysis Intelligence Index v4.3.x, with an estimated cost of $1.13 per task in the comparison. The author says GPT-6.1 Sol at medium, high and xhigh settings, GLM-5.3-Flash, DeepSeek V4.1 Flash, Claude Opus 5.5 at low, and Claude Sonnet 5.5 at medium all score higher while costing less per task. The reported alternatives range from $0.21 to $0.59 per task.
The author says Mistral’s two-week launch discount reduces the estimated cost to $0.57 per task. At that price, six of the seven configurations still cost less while scoring higher; Sonnet 5.5 at medium no longer does. The report argues that the result is more relevant to its decision than the gap between Mistral and top-scoring models, because a new model can be behind the frontier yet still offer a useful price-and-capability trade-off.
The analysis also reports that Mistral Large 4 generated 200 million output tokens across the benchmark, compared with a median of 81 million for comparable models and 25 million for GPT-6.1 Sol at high. The author says this output volume offsets Mistral’s relatively low per-token prices. These figures describe the source’s benchmark run and cost methodology; they do not establish what a particular customer would pay for a different task.
Opus builds, Sol digs, Jev decides — and Mistral Large 4 doesn’t make the cut
Two models landed on the price curve in eight days. My stack doesn’t change. The test is the same as in September: does it clear my bar at a lower cost per task than what I already run? For Mistral Large 4, no — seven configurations beat it on score and cost at once.
Cheaper per token ($4.18 vs Sol’s $10 output) — but ~8× the output of Sol for a lower score. Budget cost per task, not per token.
Speed: 116 tok/s, 1.46 s to first token — far faster than Sol at high/xhigh (57–69 s).
Cyber: AA Cyber Index 50, CyberGym-E2E 82%.
Jurisdiction: French parent, weights at the end of October.
For legally bound buyers, the best European option by a wide margin. For my stack, none of it clears the bar.
Being behind the frontier is normal for a challenger. Being beaten on both axes by models you can already buy is a pricing problem, not a capability one. Argon is the more interesting arrival — level with Astra and Fable at a third of Fable’s cost — but access is still restricted, and a model I can’t put into production isn’t part of my stack. Run the dominance test on every new model before you read its benchmark table. Most new models don’t change anything — the ruler tells you which ones do.
How the Cost Test Shapes Model Choice
The comparison highlights a practical purchasing question: cost per completed task can matter more than the price of each token. A model with inexpensive output tokens can still be costly if it produces far more text to achieve a result. For teams running repeated or multi-step jobs, that difference may affect budgets as well as speed and reliability.
The author says the Intelligence Index includes agent-oriented benchmarks, including AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0. The report interprets Mistral’s lower score as a concern for long-running tasks, where errors or extra output can accumulate across steps. That is the author’s assessment, not a finding that the model will fail on every agentic workload. Organizations would need to test their own tasks before drawing that conclusion.
The report does identify reasons some users might still consider Mistral Large 4. It cites 116 tokens per second, a 1.46-second time to first token, a score of 50 on the AA Cyber Index and an 82% result on CyberGym-E2E. The author also says model weights are expected at the end of October under EU jurisdiction, which could make the release relevant to users with European legal or operational requirements.
As an affiliate, we earn on qualifying purchases.
A Stack Built Around Task Roles
The article follows an earlier argument by its author that AI model selection had become a matter of comparing price and capability, rather than simply ranking models on a leaderboard. The new report adds Gemini 4 Argon, which the source says arrived on September 30, and Mistral Large 4, described as a recent preview release. It also includes Chinese open-weight models as reference points.
The author’s current stack assigns different jobs to different systems: Opus 5.5 handles building, GPT-6.1 Sol reviews details and changes, and Jev handles high-volume yes-or-no decisions and routing. The report says the arrival of Mistral Large 4 does not change that arrangement. It also identifies Gemini 4 Argon as a possible second-opinion model, but says access remains restricted and general availability is pending.
The source cautions that Artificial Analysis Intelligence Index v4.3.x is a map of general capability, not a verdict on an individual user’s workload. Its cost-per-task figures are presented as part of the author’s comparison. The report recommends shadow-testing before switching models, which means evaluating a candidate alongside an existing system before relying on it for production work.
“The answer is no.”
— ThorstenMeyerAI.com author
Limits of the Benchmark Comparison
The source does not provide enough detail here to independently reproduce every score or task-cost estimate, including the precise prompts, pricing assumptions and calculation method behind the per-task figures. The results should be read as the author’s comparison, not as an independently verified ranking or a guarantee of relative performance on commercial workloads.
The author’s observations of confident hallucinations are explicitly presented as personal testing, not as Artificial Analysis benchmark data. The report does not give a detailed test protocol or sample size for those observations. Nor does the general index establish how models compare on a specific company’s codebase, documents or workflows.
Availability and release details may also change. The source describes Gemini 4 Argon access as restricted and Mistral Large 4 as a preview, and says Mistral’s weights are due later in October. It does not establish whether those access conditions or release plans have since changed. Mistral’s performance under the cited introductory discount may also differ from its regular pricing.
Testing Releases Before Adoption
The author says the current model assignments will remain in place and plans to test Gemini 4 Argon if it becomes publicly available. The report gives no confirmed date for broader access. For Mistral Large 4, the next development identified is the expected release of weights at the end of October; the source does not confirm that release has occurred.
For readers weighing a change, the article’s proposed next step is to shadow-test candidates on real tasks, comparing output quality, completion time and total cost rather than relying on a general benchmark alone. Any adoption decision will depend on a model’s performance against the user’s own requirements, including any need for European jurisdiction, speed or particular specialist capabilities.
Key Questions
Why did the author reject Mistral Large 4?
The author says it scored 38 on the cited index and cost an estimated $1.13 per task, while seven configurations scored higher and cost less in the comparison.
Does the report prove Mistral Large 4 is a poor choice for everyone?
No. The source says the index measures general capability rather than performance on every workload. Its conclusion applies to the author’s stated cost-and-quality criteria; users should test their own tasks.
What strengths does the report attribute to Mistral Large 4?
The author cites 116 tokens per second, a 1.46-second time to first token, a score of 50 on the AA Cyber Index and an 82% CyberGym-E2E result. These are figures and claims reported by the source.
What does the report say about Gemini 4 Argon?
It reports a score of 53 and an estimated cost of about $1.99 per task at introductory pricing, but says access was restricted and general availability was pending. The source says the author will test it if it becomes public.
What should organizations do before switching models?
The report recommends shadow-testing: compare a candidate against the current model on representative work, tracking quality and total task costs before adopting it.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
