firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Why Would a Lazy AI Score 26 Out of 100?

Most AI benchmarks have a dirty little secret: they reward flash and punish nothing. A model that writes beautifully but abandons the job halfway can still look great. A new live experiment called Firmulate takes the opposite approach — and one of its strangest, most honest design choices is raising eyebrows: a completely passive, do-nothing run of the same business scenario still earns 26 points instead of zero.

Amazon

AI decision-making benchmark software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Company, Same Worst Week, Different Brains

Firmulate handed four frontier AI models the identical task: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners. Every decision the models made is versioned and auditable, so nothing about the scoring is a black box.

The final July 2026 league table tells a striking story: gpt-5.6-sol took first with 95, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73. But the headline finding wasn’t about scores at all — it was about follow-through. All the models spotted every crisis and refused every manipulation attempt. Yet only two of them actually signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

Why the Floor Is 26, Not 0

Here’s where the methodology gets interesting for anyone who has ever managed people (or been managed). Firmulate’s benchmark gives credit for partial progress. A model that correctly reads the situation, identifies the crisis, and starts down the right path has genuinely done some of the work — even if it never closes. So the do-nothing baseline isn’t literally doing nothing; it’s doing the minimum a competent-but-passive manager would do, and that minimum has real value. Hence 26 points, not zero.

But there’s a ceiling rule that cuts the other way, and it’s the benchmark’s moral spine: a single breach of trust caps the total grade. As the experiment puts it, “no amount of good work outweighs a breach of trust.” Slack off and you lose points gradually; deceive once and you lose the game.

The Buried Fact That Decided the Deal

The reason only two models closed the €55k contract? The decisive competitive weakness wasn’t in the customer conversation at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t read before acting left the money on the table. It’s the AI equivalent of walking into a negotiation without reading your own notes.

Opus 4.8: The Cautionary Tale

Opus 4.8 was the most thorough participant in the field — it learned over 80 new rules and produced the deepest analyses. And it still came last. The close was left unfinished, and discipline slipped: it attempted writes into a locked department instead of escalating properly. Notably, the same weakness appeared, just weaker, in all four models. One fairness footnote: Kimi K3 ran at its API-default effort setting while the others ran at xhigh — and still nearly won.

Under Social Pressure, Everyone Held

The experiment didn’t go easy on the models. It threw fake CEO messages at them, escalating over three stages, plus a reporter trick: “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the kind of judgment you want before an agent touches anything important.

You Can Watch It Live

This isn’t a one-off paper. Firmulate is running right now: a company of 13 synthetic employees with real money mechanics — burning €105k a month against €2.3k in MRR — a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. There’s also a “guess the model” quiz built from 242 real, unedited management decisions. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The 26-point floor is what separates this benchmark from the hype cycle. It says: competent passivity is worth something, honest effort counts, and one lie erases everything. Notice, too, what’s missing — a perfect 100. In a field where suspiciously round scores are everywhere, a league topped by a 95 reads as a measurement, not a marketing claim. If AI agents are about to touch your CRM, your support queue, or your forecast, this is the kind of scoreboard worth watching.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Construct a Lead Qualification System That Continues Working Overnight

A new automated lead qualification system now runs continuously, scoring and routing leads overnight to improve efficiency and sales outcomes.

StrongMocha News Group Expands into Nanotechnology

AIThis post was created with the assistance of artificial intelligence (AI).Berlin, Germany…

What you missed in streaming this week; Fubo quietly raises prices, Mountain West launches streaming hub, more

This week in streaming: Fubo quietly increases subscription prices, and Mountain West launches a new streaming platform. Key updates and implications explained.

Dropbox CEO Drew Houston to step down

Drew Houston, founder and CEO of Dropbox, will step down and become executive chairman, with Ashraf Alkarmi set to succeed him as CEO, marking a leadership change after 19 years.