Opus Builds, Sol Digs, Jev Decides — And Mistral Large 4 Doesn’t Make The Cut: My October 2026 AI Stack
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Opus Builds, Sol Digs, Jev Decides — And Mistral Large 4 Doesn’t Make The Cut: My October 2026 AI Stack on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

ThorstenMeyerAI.com reports that Mistral Large 4 did not meet the author’s cost-and-capability threshold in a comparison using Artificial Analysis Intelligence Index v4.3.x scores and estimated task costs. The author says seven model configurations score higher and cost less per task, while noting Mistral’s speed, cyber benchmarks and expected European weights as potential advantages. These are one developer’s findings, not a universal assessment of model performance.

ThorstenMeyerAI.com reports that its author will not add Mistral Large 4 to a working AI model stack after comparing it with other systems on benchmark scores and estimated cost per task. In the author’s analysis, seven configurations score higher and cost less, a finding relevant to developers choosing models by both capability and operating expense. The comparison relies on a general benchmark and the author’s cost estimates, not tests of every real-world workload.

The report places Mistral Large 4 at 38 points on the Artificial Analysis Intelligence Index v4.3.x, with an estimated cost of $1.13 per task in the comparison. The author says GPT-6.1 Sol at medium, high and xhigh settings, GLM-5.3-Flash, DeepSeek V4.1 Flash, Claude Opus 5.5 at low, and Claude Sonnet 5.5 at medium all score higher while costing less per task. The reported alternatives range from $0.21 to $0.59 per task.

The author says Mistral’s two-week launch discount reduces the estimated cost to $0.57 per task. At that price, six of the seven configurations still cost less while scoring higher; Sonnet 5.5 at medium no longer does. The report argues that the result is more relevant to its decision than the gap between Mistral and top-scoring models, because a new model can be behind the frontier yet still offer a useful price-and-capability trade-off.

The analysis also reports that Mistral Large 4 generated 200 million output tokens across the benchmark, compared with a median of 81 million for comparable models and 25 million for GPT-6.1 Sol at high. The author says this output volume offsets Mistral’s relatively low per-token prices. These figures describe the source’s benchmark run and cost methodology; they do not establish what a particular customer would pay for a different task.

At a glance
reportWhen: Published after Mistral Large 4’s repor…
The developmentAn AI developer published a comparison concluding that Mistral Large 4 is outperformed on both benchmark score and estimated task cost by seven configurations already available to the author.
My October 2026 AI Stack — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Opus builds, Sol digs, Jev decides — and Mistral Large 4 doesn’t make the cut

Two models landed on the price curve in eight days. My stack doesn’t change. The test is the same as in September: does it clear my bar at a lower cost per task than what I already run? For Mistral Large 4, no — seven configurations beat it on score and cost at once.

Builds
Opus 5.5
high · xhigh for hard problems
Digs & reviews
GPT-6.1 Sol
high or xhigh · $0.32–0.39
Decides
Jev
high-volume yes/no & routing
Evaluated · not adopted
Mistral Large 4
dominated on score and cost
The dominance test — everything in the green box beats it on both axes
BETTER AND CHEAPER THAN MISTRAL LARGE 4$0.05$0.10$0.50$1$5$1030354045505560cost per Intelligence Index task (log scale) → cheaper to the leftindex ↑Opus 5.5Sonnet 5.5GPT-6.1 SolFable 5.1AstraArgon*LunaGLM-5.3-FlashDeepSeek V4.1 FlashMistral Large 4 · 38 · $1.13promo $0.57
Artificial Analysis Intelligence Index v4.3.x. Lines show effort settings. *Argon: restricted access, introductory price (~$1.99). Hollow orange dot: Mistral’s two-week launch discount — six of the seven still dominate at that price.
Seven configurations better and cheaper than Mistral Large 4 (38 · $1.13)
Configuration
Index
$ / task
vs Mistral Large 4
GPT-6.1 Sol · xhigh
51
$0.39
+13 pts · 2.9× cheaper
GPT-6.1 Sol · high
50
$0.32
+12 pts · 3.5× cheaper
GPT-6.1 Sol · medium
48
$0.21
+10 pts · 5.4× cheaper
GLM-5.3-Flash · open
42
$0.25
+4 pts · 4.5× cheaper
Claude Opus 5.5 · low
42
$0.55
+4 pts · 2.1× cheaper
Claude Sonnet 5.5 · medium
41
$0.59
+3 pts · 1.9× cheaper
DeepSeek V4.1 Flash · open
39
$0.27
+1 pt · 4.2× cheaper
Mistral Large 4 · preview
38
$1.13
$0.57 at launch discount
Why so expensive: output tokens on the Index
Mistral Large 4200M
Median, comparable81M
GPT-6.1 Sol · high25M

Cheaper per token ($4.18 vs Sol’s $10 output) — but ~8× the output of Sol for a lower score. Budget cost per task, not per token.

What it still has going for it

Speed: 116 tok/s, 1.46 s to first token — far faster than Sol at high/xhigh (57–69 s).
Cyber: AA Cyber Index 50, CyberGym-E2E 82%.
Jurisdiction: French parent, weights at the end of October.
For legally bound buyers, the best European option by a wide margin. For my stack, none of it clears the bar.

The take

Being behind the frontier is normal for a challenger. Being beaten on both axes by models you can already buy is a pricing problem, not a capability one. Argon is the more interesting arrival — level with Astra and Fable at a third of Fable’s cost — but access is still restricted, and a model I can’t put into production isn’t part of my stack. Run the dominance test on every new model before you read its benchmark table. Most new models don’t change anything — the ruler tells you which ones do.

Sources: Artificial Analysis Intelligence Index v4.3.x — Mistral Large 4 Preview article & model page (6 Oct 2026); GLM-5.3-Flash and DeepSeek V4.1 Flash per AA; Gemini 4 Argon via AA-derived reporting; all other scores and costs from “Opus Builds, Sol Digs, Jev Decides: My September 2026 AI Stack” (29 Sep 2026). Dominance ratios are the author’s arithmetic. Not investment advice.
thorstenmeyerai.com

How the Cost Test Shapes Model Choice

The comparison highlights a practical purchasing question: cost per completed task can matter more than the price of each token. A model with inexpensive output tokens can still be costly if it produces far more text to achieve a result. For teams running repeated or multi-step jobs, that difference may affect budgets as well as speed and reliability.

The author says the Intelligence Index includes agent-oriented benchmarks, including AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0. The report interprets Mistral’s lower score as a concern for long-running tasks, where errors or extra output can accumulate across steps. That is the author’s assessment, not a finding that the model will fail on every agentic workload. Organizations would need to test their own tasks before drawing that conclusion.

The report does identify reasons some users might still consider Mistral Large 4. It cites 116 tokens per second, a 1.46-second time to first token, a score of 50 on the AA Cyber Index and an 82% result on CyberGym-E2E. The author also says model weights are expected at the end of October under EU jurisdiction, which could make the release relevant to users with European legal or operational requirements.

Amazon

AI model benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Stack Built Around Task Roles

The article follows an earlier argument by its author that AI model selection had become a matter of comparing price and capability, rather than simply ranking models on a leaderboard. The new report adds Gemini 4 Argon, which the source says arrived on September 30, and Mistral Large 4, described as a recent preview release. It also includes Chinese open-weight models as reference points.

The author’s current stack assigns different jobs to different systems: Opus 5.5 handles building, GPT-6.1 Sol reviews details and changes, and Jev handles high-volume yes-or-no decisions and routing. The report says the arrival of Mistral Large 4 does not change that arrangement. It also identifies Gemini 4 Argon as a possible second-opinion model, but says access remains restricted and general availability is pending.

The source cautions that Artificial Analysis Intelligence Index v4.3.x is a map of general capability, not a verdict on an individual user’s workload. Its cost-per-task figures are presented as part of the author’s comparison. The report recommends shadow-testing before switching models, which means evaluating a candidate alongside an existing system before relying on it for production work.

“The answer is no.”

— ThorstenMeyerAI.com author

Limits of the Benchmark Comparison

The source does not provide enough detail here to independently reproduce every score or task-cost estimate, including the precise prompts, pricing assumptions and calculation method behind the per-task figures. The results should be read as the author’s comparison, not as an independently verified ranking or a guarantee of relative performance on commercial workloads.

The author’s observations of confident hallucinations are explicitly presented as personal testing, not as Artificial Analysis benchmark data. The report does not give a detailed test protocol or sample size for those observations. Nor does the general index establish how models compare on a specific company’s codebase, documents or workflows.

Availability and release details may also change. The source describes Gemini 4 Argon access as restricted and Mistral Large 4 as a preview, and says Mistral’s weights are due later in October. It does not establish whether those access conditions or release plans have since changed. Mistral’s performance under the cited introductory discount may also differ from its regular pricing.

Testing Releases Before Adoption

The author says the current model assignments will remain in place and plans to test Gemini 4 Argon if it becomes publicly available. The report gives no confirmed date for broader access. For Mistral Large 4, the next development identified is the expected release of weights at the end of October; the source does not confirm that release has occurred.

For readers weighing a change, the article’s proposed next step is to shadow-test candidates on real tasks, comparing output quality, completion time and total cost rather than relying on a general benchmark alone. Any adoption decision will depend on a model’s performance against the user’s own requirements, including any need for European jurisdiction, speed or particular specialist capabilities.

Key Questions

Why did the author reject Mistral Large 4?

The author says it scored 38 on the cited index and cost an estimated $1.13 per task, while seven configurations scored higher and cost less in the comparison.

Does the report prove Mistral Large 4 is a poor choice for everyone?

No. The source says the index measures general capability rather than performance on every workload. Its conclusion applies to the author’s stated cost-and-quality criteria; users should test their own tasks.

What strengths does the report attribute to Mistral Large 4?

The author cites 116 tokens per second, a 1.46-second time to first token, a score of 50 on the AA Cyber Index and an 82% CyberGym-E2E result. These are figures and claims reported by the source.

What does the report say about Gemini 4 Argon?

It reports a score of 53 and an estimated cost of about $1.99 per task at introductory pricing, but says access was restricted and general availability was pending. The source says the author will test it if it becomes public.

What should organizations do before switching models?

The report recommends shadow-testing: compare a candidate against the current model on representative work, tracking quality and total task costs before adopting it.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Design Maven Lisa Marie Holmes Elevates Renovation

Nurture your renovation dreams with design maven Lisa Marie Holmes as she unveils her transformative expertise and innovative solutions.

Europe Regulated the Interface and Forgot to Build the Engine

Europe focused on regulating AI interfaces like cookie banners but has failed to build the underlying AI technology, falling behind global leaders.

India’s Tata Sons faces growing IPO pressure after RBI rule change

India’s Tata Sons is under growing pressure to list its shares following RBI’s revised classification of shadow lenders, raising questions on corporate transparency.

RJ Scaringe has raised more than $12 billion across three startups and investors still want more

Serial entrepreneur RJ Scaringe has raised more than $12 billion for his three startups, including Rivian, Also, and Mind Robotics, highlighting strong investor confidence.