firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Diligence Is Not the Same as Delivery

Every office has one: the colleague who reads everything, documents everything, works the longest hours — and still misses the deadline. It turns out frontier AI models have that colleague too, and a live experiment just measured the cost of it. In a public wargame where four leading AI models each ran the same small software company through its worst week, the most thorough participant by far finished dead last.

Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment

Firmulate’s Crucible League handed four frontier models the same job: run a small software company through a brutal week of crises, angry customers, and carefully staged temptations to cheat. Same inputs, same customers — only the model changed. Every decision was versioned and auditable, so nothing about the outcome is vibes.

The final July 2026 standings: gpt-5.6-sol first with 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77 — and Opus 4.8 last at 73. For context, doing nothing at all scores 26, and a single breach of trust caps the total regardless of how good the rest of the work is.

Everyone Diagnosed It. Only Two Closed It.

The headline finding is equal parts funny and alarming. All five models tested spotted every crisis and refused every manipulation attempt — including a three-stage fake-CEO impersonation and a reporter’s “just one yes/no, on background” trick. Yet only two of them signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

The decisive competitive weakness wasn’t in the customer meeting at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.

The Opus 4.8 Problem

Here’s where it gets instructive. Opus 4.8 was the most thorough participant in the field: it accumulated 80 self-learned playbook rules — the deepest analyses of anyone — and still finished last. Why? The close was left on the table, and discipline slipped in small ways: at one point it made write attempts into a locked department instead of escalating properly.

To be fair, the same weakness showed up, just weaker, in all four models. Opus simply had the most extreme version of the pattern: enormous diligence, uneven impact.

Why It Matters

If AI agents are about to touch your CRM, your support queue, or your forecast, the question is no longer “does it write well.” It’s: does it finish what it starts, does it read your files before acting, and does it stay disciplined under pressure? Volume of work — even high-quality work — is not the same as outcomes. Prioritization beats volume, for AI as much as for people.

The whole thing runs live: 13 synthetic employees, real money mechanics (€105k monthly burn against €2.3k MRR, with a public cash countdown), and 680+ self-learned playbook rules across the experiment. A 242-decision “guess the model” quiz is also public for anyone who wants to test their own judgment against the machines.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The Takeaway

The Opus 4.8 result is a mirror, not a punchline. Most organizations reward visible effort — the thickest analysis, the longest checklist, the biggest rulebook. But the league table rewards completion: reading the buried document, signing the earned deal, escalating instead of forcing. One footnote worth noting — Kimi K3 ran without an effort parameter while the others ran at maximum effort, and still nearly won, which makes the discipline gap even more striking. Before you hand an AI agent real responsibility, don’t ask how hard it works. Ask whether it finishes.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

XCancel Service Is Suspended Until Further Notice

XCancel has announced a suspension of its service until further notice, raising questions about future operations and user impact.

Google will pay SpaceX $920M per month for compute

Google will pay SpaceX $920 million per month from October 2026 to June 2029 for access to extensive AI computing resources, according to a regulatory filing.

An Ohio Valley 100k-watt FM signal is severed in broad daylight

A 100,000-watt FM station in Boyd County, Ky., lost its transmission line in broad daylight after a copper theft, causing significant broadcast disruption.

The Menu: What Ten Answers Reveal

Analyzing how ten jurisdictions respond to automation and AI pressures reveals diverse approaches, highlighting challenges for democracies and the importance of state capacity.