firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Why Would a Lazy AI Score 26 Out of 100?

Most AI benchmarks have a dirty little secret: they reward flash and punish nothing. A model that writes beautifully but abandons the job halfway can still look great. A new live experiment called Firmulate takes the opposite approach — and one of its strangest, most honest design choices is raising eyebrows: a completely passive, do-nothing run of the same business scenario still earns 26 points instead of zero.

Amazon

AI decision-making benchmark software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Company, Same Worst Week, Different Brains

Firmulate handed four frontier AI models the identical task: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners. Every decision the models made is versioned and auditable, so nothing about the scoring is a black box.

The final July 2026 league table tells a striking story: gpt-5.6-sol took first with 95, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73. But the headline finding wasn’t about scores at all — it was about follow-through. All the models spotted every crisis and refused every manipulation attempt. Yet only two of them actually signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

Why the Floor Is 26, Not 0

Here’s where the methodology gets interesting for anyone who has ever managed people (or been managed). Firmulate’s benchmark gives credit for partial progress. A model that correctly reads the situation, identifies the crisis, and starts down the right path has genuinely done some of the work — even if it never closes. So the do-nothing baseline isn’t literally doing nothing; it’s doing the minimum a competent-but-passive manager would do, and that minimum has real value. Hence 26 points, not zero.

But there’s a ceiling rule that cuts the other way, and it’s the benchmark’s moral spine: a single breach of trust caps the total grade. As the experiment puts it, “no amount of good work outweighs a breach of trust.” Slack off and you lose points gradually; deceive once and you lose the game.

The Buried Fact That Decided the Deal

The reason only two models closed the €55k contract? The decisive competitive weakness wasn’t in the customer conversation at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t read before acting left the money on the table. It’s the AI equivalent of walking into a negotiation without reading your own notes.

Opus 4.8: The Cautionary Tale

Opus 4.8 was the most thorough participant in the field — it learned over 80 new rules and produced the deepest analyses. And it still came last. The close was left unfinished, and discipline slipped: it attempted writes into a locked department instead of escalating properly. Notably, the same weakness appeared, just weaker, in all four models. One fairness footnote: Kimi K3 ran at its API-default effort setting while the others ran at xhigh — and still nearly won.

Under Social Pressure, Everyone Held

The experiment didn’t go easy on the models. It threw fake CEO messages at them, escalating over three stages, plus a reporter trick: “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the kind of judgment you want before an agent touches anything important.

You Can Watch It Live

This isn’t a one-off paper. Firmulate is running right now: a company of 13 synthetic employees with real money mechanics — burning €105k a month against €2.3k in MRR — a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. There’s also a “guess the model” quiz built from 242 real, unedited management decisions. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The 26-point floor is what separates this benchmark from the hype cycle. It says: competent passivity is worth something, honest effort counts, and one lie erases everything. Notice, too, what’s missing — a perfect 100. In a field where suspiciously round scores are everywhere, a league topped by a 95 reads as a measurement, not a marketing claim. If AI agents are about to touch your CRM, your support queue, or your forecast, this is the kind of scoreboard worth watching.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Kv4p HT – A homebrew 1W radio (VHF or UHF) that plugs into an Android phone

Open-source kv4p HT transforms Android phones into 1W VHF/UHF ham transceivers, requiring DIY assembly and amateur radio licensing. Details inside.

Age verification tech could put children at greater risk, says think tank

A UK think tank warns that mandatory online age verification could expose children to greater risks, including privacy breaches and marginalization.

Several Overwatch heroes are about to hit Fortnite

Epic Games confirms several Overwatch characters will join Fortnite as skins on May 14, including Mercy, Tracer, Genji, and D.Va, with additional cosmetics and features planned.

Uv is fantastic, but its package management UX is a mess

Uv is praised for its speed and Python handling but criticized for poor package management UX, including unsafe defaults and clunky commands.