
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Leaderboard Won’t Tell You If Your AI Agent Quits at the Finish Line
Every few months, a new AI model tops a coding benchmark or chat arena, and the internet declares a winner. But those tests measure something narrow: how well a model answers a question when there’s no clock, no cash running out, and nobody trying to trick it. A live experiment at Firmulate just measured something different — what happens when frontier AI models are handed a real, money-losing company and told to run it through its worst week. The results expose a gap that no leaderboard currently captures.
AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Four Models, One Terrible Week
The setup is elegantly cruel. Four frontier AI models were each given the same job: run the same small software company through an identical gauntlet — a churn wave, a price increase, a downround scenario, a PR crisis, and a stream of social-engineering attempts. Same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable.
The Crucible League’s final July 2026 standings: gpt-5.6-sol scored 95, Kimi K3 scored 93, Sonnet 5 scored 88, Fable 5 scored 77, and Opus 4.8 scored 73. A do-nothing baseline scored 26 — partial progress counts, but a single breach of trust caps the total. As the rules put it: “no amount of good work outweighs a breach of trust.”
Everyone Diagnosed the Disease. Two Prescribed the Cure.
The headline finding is strange and important: all four models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. Imagine a consultant who nails the diagnosis, delivers a flawless presentation, and then simply… forgets to ask the client to sign.
The deal turned on a buried fact. The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer event itself. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t.
The Social Engineering Test: 5 for 5
Then came the pressure. Fake CEO messages escalated over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. All five participating models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the right instinct — and notably, K3 achieved its second-place score while running without an effort parameter (API default), while the others ran at xhigh.
The Thoroughness Paradox
Opus 4.8 is the cautionary tale of the batch. It was the most thorough participant — over 80 learned rules, the deepest analyses — yet finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Knowing a lot, it turns out, is not the same as finishing what you start.
You Can Watch the Company Lose Money in Real Time
Firmulate isn’t a slide deck. The live company has 13 synthetic employees, real money mechanics — burning €105k per month against €2.3k in MRR — with a public cash countdown, over 680 self-learned playbook rules, and every workday versioned. It runs every business day and is watchable at firmulate.com. The site rebuilds itself twice a day, currently at company day 1131.
Want to test your own instincts? A quiz powered by 242 real, unedited management decisions lets you guess which model made which call. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Full results and plain-language findings are on the benchmarks page.

Management Quality, Not Chat Quality
If AI agents will soon touch your CRM, your support queue, or your forecast, the real question isn’t “does it write well?” It’s: does it finish what it starts, does it read your files before acting, does it stay honest under pressure — and what does a unit of useful work actually cost? The Firmulate experiment suggests those are different skills, measured differently, and that today’s leaderboards mostly measure the wrong one. The model that wins the arena isn’t necessarily the model that closes the deal.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Back to school Picks
back to school
As an affiliate, we earn on qualifying purchases.