
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
Diligence Is Not the Same as Delivery
Every office has one: the colleague who reads everything, documents everything, works the longest hours — and still misses the deadline. It turns out frontier AI models have that colleague too, and a live experiment just measured the cost of it. In a public wargame where four leading AI models each ran the same small software company through its worst week, the most thorough participant by far finished dead last.
As an affiliate, we earn on qualifying purchases.
The Experiment
Firmulate’s Crucible League handed four frontier models the same job: run a small software company through a brutal week of crises, angry customers, and carefully staged temptations to cheat. Same inputs, same customers — only the model changed. Every decision was versioned and auditable, so nothing about the outcome is vibes.
The final July 2026 standings: gpt-5.6-sol first with 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77 — and Opus 4.8 last at 73. For context, doing nothing at all scores 26, and a single breach of trust caps the total regardless of how good the rest of the work is.
Everyone Diagnosed It. Only Two Closed It.
The headline finding is equal parts funny and alarming. All five models tested spotted every crisis and refused every manipulation attempt — including a three-stage fake-CEO impersonation and a reporter’s “just one yes/no, on background” trick. Yet only two of them signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
The decisive competitive weakness wasn’t in the customer meeting at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.
The Opus 4.8 Problem
Here’s where it gets instructive. Opus 4.8 was the most thorough participant in the field: it accumulated 80 self-learned playbook rules — the deepest analyses of anyone — and still finished last. Why? The close was left on the table, and discipline slipped in small ways: at one point it made write attempts into a locked department instead of escalating properly.
To be fair, the same weakness showed up, just weaker, in all four models. Opus simply had the most extreme version of the pattern: enormous diligence, uneven impact.
Why It Matters
If AI agents are about to touch your CRM, your support queue, or your forecast, the question is no longer “does it write well.” It’s: does it finish what it starts, does it read your files before acting, and does it stay disciplined under pressure? Volume of work — even high-quality work — is not the same as outcomes. Prioritization beats volume, for AI as much as for people.
The whole thing runs live: 13 synthetic employees, real money mechanics (€105k monthly burn against €2.3k MRR, with a public cash countdown), and 680+ self-learned playbook rules across the experiment. A 242-decision “guess the model” quiz is also public for anyone who wants to test their own judgment against the machines.

The Takeaway
The Opus 4.8 result is a mirror, not a punchline. Most organizations reward visible effort — the thickest analysis, the longest checklist, the biggest rulebook. But the league table rewards completion: reading the buried document, signing the earned deal, escalating instead of forcing. One footnote worth noting — Kimi K3 ran without an effort parameter while the others ran at maximum effort, and still nearly won, which makes the discipline gap even more striking. Before you hand an AI agent real responsibility, don’t ask how hard it works. Ask whether it finishes.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.