firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

The Leaderboard Won’t Tell You If Your AI Agent Quits at the Finish Line

Every few months, a new AI model tops a coding benchmark or chat arena, and the internet declares a winner. But those tests measure something narrow: how well a model answers a question when there’s no clock, no cash running out, and nobody trying to trick it. A live experiment at Firmulate just measured something different — what happens when frontier AI models are handed a real, money-losing company and told to run it through its worst week. The results expose a gap that no leaderboard currently captures.

Amazon

AI decision-making software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Four Models, One Terrible Week

The setup is elegantly cruel. Four frontier AI models were each given the same job: run the same small software company through an identical gauntlet — a churn wave, a price increase, a downround scenario, a PR crisis, and a stream of social-engineering attempts. Same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable.

The Crucible League’s final July 2026 standings: gpt-5.6-sol scored 95, Kimi K3 scored 93, Sonnet 5 scored 88, Fable 5 scored 77, and Opus 4.8 scored 73. A do-nothing baseline scored 26 — partial progress counts, but a single breach of trust caps the total. As the rules put it: “no amount of good work outweighs a breach of trust.”

Everyone Diagnosed the Disease. Two Prescribed the Cure.

The headline finding is strange and important: all four models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. Imagine a consultant who nails the diagnosis, delivers a flawless presentation, and then simply… forgets to ask the client to sign.

The deal turned on a buried fact. The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer event itself. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t.

The Social Engineering Test: 5 for 5

Then came the pressure. Fake CEO messages escalated over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. All five participating models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the right instinct — and notably, K3 achieved its second-place score while running without an effort parameter (API default), while the others ran at xhigh.

The Thoroughness Paradox

Opus 4.8 is the cautionary tale of the batch. It was the most thorough participant — over 80 learned rules, the deepest analyses — yet finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Knowing a lot, it turns out, is not the same as finishing what you start.

You Can Watch the Company Lose Money in Real Time

Firmulate isn’t a slide deck. The live company has 13 synthetic employees, real money mechanics — burning €105k per month against €2.3k in MRR — with a public cash countdown, over 680 self-learned playbook rules, and every workday versioned. It runs every business day and is watchable at firmulate.com. The site rebuilds itself twice a day, currently at company day 1131.

Want to test your own instincts? A quiz powered by 242 real, unedited management decisions lets you guess which model made which call. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Full results and plain-language findings are on the benchmarks page.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

Management Quality, Not Chat Quality

If AI agents will soon touch your CRM, your support queue, or your forecast, the real question isn’t “does it write well?” It’s: does it finish what it starts, does it read your files before acting, does it stay honest under pressure — and what does a unit of useful work actually cost? The Firmulate experiment suggests those are different skills, measured differently, and that today’s leaderboards mostly measure the wrong one. The model that wins the arena isn’t necessarily the model that closes the deal.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


BACK TO SCHOOL

Back to school Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Global infrastructure funding doubles over 5 years, led by Japanese banks

Worldwide infrastructure project financing has doubled over five years, with Japanese banks, especially MUFG, at the forefront, driven by supply chain diversification and geopolitical risks.

All Vehicles Sold in the EU Must Be Able to Hook Up to a Breathalyzer

Starting July 1, all vehicles sold in the EU must include an interface for installing breathalyzer ignition locks, part of efforts to reduce drunk-driving fatalities.

NYT Connections today – my hints and answers for June 30 (#1115)

Complete hints and solutions for NYT Connections puzzle #1115 on June 30, including tips and key details.

Silicon Valley’s vacationland needs a new energy provider just as AI is driving prices up

Lake Tahoe’s power provider contract ends in May 2027 amid rising energy demands from AI data centers, prompting regional supply concerns.