firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

Diligence Is Not the Same as Delivery

Every office has one: the colleague who reads everything, documents everything, works the longest hours — and still misses the deadline. It turns out frontier AI models have that colleague too, and a live experiment just measured the cost of it. In a public wargame where four leading AI models each ran the same small software company through its worst week, the most thorough participant by far finished dead last.

Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment

Firmulate’s Crucible League handed four frontier models the same job: run a small software company through a brutal week of crises, angry customers, and carefully staged temptations to cheat. Same inputs, same customers — only the model changed. Every decision was versioned and auditable, so nothing about the outcome is vibes.

The final July 2026 standings: gpt-5.6-sol first with 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77 — and Opus 4.8 last at 73. For context, doing nothing at all scores 26, and a single breach of trust caps the total regardless of how good the rest of the work is.

Everyone Diagnosed It. Only Two Closed It.

The headline finding is equal parts funny and alarming. All five models tested spotted every crisis and refused every manipulation attempt — including a three-stage fake-CEO impersonation and a reporter’s “just one yes/no, on background” trick. Yet only two of them signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

The decisive competitive weakness wasn’t in the customer meeting at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.

The Opus 4.8 Problem

Here’s where it gets instructive. Opus 4.8 was the most thorough participant in the field: it accumulated 80 self-learned playbook rules — the deepest analyses of anyone — and still finished last. Why? The close was left on the table, and discipline slipped in small ways: at one point it made write attempts into a locked department instead of escalating properly.

To be fair, the same weakness showed up, just weaker, in all four models. Opus simply had the most extreme version of the pattern: enormous diligence, uneven impact.

Why It Matters

If AI agents are about to touch your CRM, your support queue, or your forecast, the question is no longer “does it write well.” It’s: does it finish what it starts, does it read your files before acting, and does it stay disciplined under pressure? Volume of work — even high-quality work — is not the same as outcomes. Prioritization beats volume, for AI as much as for people.

The whole thing runs live: 13 synthetic employees, real money mechanics (€105k monthly burn against €2.3k MRR, with a public cash countdown), and 680+ self-learned playbook rules across the experiment. A 242-decision “guess the model” quiz is also public for anyone who wants to test their own judgment against the machines.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The Takeaway

The Opus 4.8 result is a mirror, not a punchline. Most organizations reward visible effort — the thickest analysis, the longest checklist, the biggest rulebook. But the league table rewards completion: reading the buried document, signing the earned deal, escalating instead of forcing. One footnote worth noting — Kimi K3 ran without an effort parameter while the others ran at maximum effort, and still nearly won, which makes the discipline gap even more striking. Before you hand an AI agent real responsibility, don’t ask how hard it works. Ask whether it finishes.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Gaming Signal Monitor: Minecraft: Java Edition Now Uses SDL3

Minecraft: Java Edition now uses SDL3 for graphics rendering, marking a significant update. Details on impact and future steps remain limited.

Southern Icons' Birthplaces Shaped Music and Culture

Immerse yourself in the profound impact of Southern birthplaces on music and culture, shaping icons whose roots continue to inspire creativity and innovation.

Saying Goodbye to Asm.js

Firefox 148 disables asm.js optimizations by default and plans to remove it entirely, encouraging developers to migrate to WebAssembly for better performance.

The Anthropic-Blackstone-Goldman JV: Reverse-Engineering the $1.5B Enterprise AI Services Structure

Anthropic, Blackstone, H&F, and Goldman Sachs form a $1.5B standalone AI services firm targeting mid-sized companies, embedding Anthropic engineers.