firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

AI model rankings usually reward answers. Firmulate puts models in charge of the same struggling software company and watches what they do when customers, cash and security are on the line. In its July 2026 Crucible league, Moonshot’s Kimi K3 finished second, beating three of four Western frontier models.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Five models, one very bad week

Firmulate gave each model the same small company, customers, crises and temptations. Every decision was versioned and auditable. The company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue. Readers can watch its public cash countdown and daily work unfold at Firmulate.

The final league table puts gpt-5.6-sol first with 95 points, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Amazon

AI management decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Reading the files made the difference

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The decisive weakness in a competitor’s offer was buried two document references deep in the company’s files, rather than in the customer event. Models that read those files won the deal at full price, worth €4,583 in monthly recurring revenue. The result captures a practical gap between diagnosing a problem and carrying the work through to a signature.

The pressure tests included fake CEO messages that escalated through three stages and a reporter asking for “just one yes/no, on background.” All five models refused. Kimi’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

K3’s second-place finish was paired with the cleanest discipline in the field: it made one deviation. Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and tried to write into a locked department instead of escalating. Firmulate says weaker versions of that discipline problem appeared in all four other participants.

A leaderboard is a starting point

For companies considering AI agents in customer support, sales or forecasting, polished chat responses are only part of the picture. Firmulate’s test asks whether a model reads the available information, handles pressure honestly and finishes consequential work. Its benchmark results offer a public comparison, while a quiz built from 242 real, unedited management decisions lets readers guess which model made each call.

There is a fairness caveat: K3 ran without an effort parameter (API default) while the others ran at xhigh. The results are therefore a useful snapshot of performance under these test conditions, not a universal ranking for every task or setup. Firmulate says enterprises can run the same wargame against a read-only export of their own business; nothing writes back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test before you hand over the keys

Kimi K3’s near-top result shows that the frontier-model leaderboard is open, while the gap between sound analysis and a completed deal shows why general rankings alone cannot settle a hiring decision. Firms weighing AI agents can use a public benchmark as a reference, then test models against the decisions their own business actually needs made.

Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Revolutionary Times: Pluto's Impact on Aquarius

Unveil the cosmic forces shaping history and stirring revolution as Pluto aligns with Aquarius, sparking profound societal upheaval and transformative change.

FiveThirtyEight articles on the Internet Archive

The Internet Archive has preserved over 21,000 pages from FiveThirtyEight, highlighting the platform’s historical data and analysis efforts.

Blanchard's Diverse Heritage and Advocacy Impact

With a blend of diverse backgrounds, Blanchard's advocacy for inclusivity in media leaves a lasting impact on Hollywood and beyond.

Exploring ChatGPT For Teens: AI That Supports Education And Safety

OpenAI announced ChatGPT for Teens, a new version aimed at supporting education with safety protections, but key details about features and rollout remain unclear.