firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

AI model rankings usually reward answers. Firmulate puts models in charge of the same struggling software company and watches what they do when customers, cash and security are on the line. In its July 2026 Crucible league, Moonshot’s Kimi K3 finished second, beating three of four Western frontier models.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Five models, one very bad week

Firmulate gave each model the same small company, customers, crises and temptations. Every decision was versioned and auditable. The company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue. Readers can watch its public cash countdown and daily work unfold at Firmulate.

The final league table puts gpt-5.6-sol first with 95 points, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Amazon

AI management decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Reading the files made the difference

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The decisive weakness in a competitor’s offer was buried two document references deep in the company’s files, rather than in the customer event. Models that read those files won the deal at full price, worth €4,583 in monthly recurring revenue. The result captures a practical gap between diagnosing a problem and carrying the work through to a signature.

The pressure tests included fake CEO messages that escalated through three stages and a reporter asking for “just one yes/no, on background.” All five models refused. Kimi’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

K3’s second-place finish was paired with the cleanest discipline in the field: it made one deviation. Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and tried to write into a locked department instead of escalating. Firmulate says weaker versions of that discipline problem appeared in all four other participants.

A leaderboard is a starting point

For companies considering AI agents in customer support, sales or forecasting, polished chat responses are only part of the picture. Firmulate’s test asks whether a model reads the available information, handles pressure honestly and finishes consequential work. Its benchmark results offer a public comparison, while a quiz built from 242 real, unedited management decisions lets readers guess which model made each call.

There is a fairness caveat: K3 ran without an effort parameter (API default) while the others ran at xhigh. The results are therefore a useful snapshot of performance under these test conditions, not a universal ranking for every task or setup. Firmulate says enterprises can run the same wargame against a read-only export of their own business; nothing writes back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test before you hand over the keys

Kimi K3’s near-top result shows that the frontier-model leaderboard is open, while the gap between sound analysis and a completed deal shows why general rankings alone cannot settle a hiring decision. Firms weighing AI agents can use a public benchmark as a reference, then test models against the decisions their own business actually needs made.

Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How To Prepare For The Future Of AI In Critical Cyber Capabilities

OpenAI releases a position paper on handling advanced AI models with cybersecurity capabilities, signaling a focus on safety and policy for critical cyber skills.

Musical Stars: Swift, Del Rey, Eilish Buzz

Prepare to be captivated by the latest buzz surrounding Taylor Swift, Lana Del Rey, and Billie Eilish in the music world, leaving you eager for more.

Postgres Rewritten In Rust, Now Passing 100% Of The Postgres Regression Tests

The Postgres database system has been fully rewritten in Rust and now passes 100% of its regression tests, marking a significant milestone in its development.

Anthropic Just Showed An Early Version Of Self-improving AI – Digital Trends

Anthropic has shown an early version of a self-improving AI, raising questions about autonomy, safety, and development speed. Details remain limited.