
AI model rankings usually reward answers. Firmulate puts models in charge of the same struggling software company and watches what they do when customers, cash and security are on the line. In its July 2026 Crucible league, Moonshot’s Kimi K3 finished second, beating three of four Western frontier models.
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Five models, one very bad week
Firmulate gave each model the same small company, customers, crises and temptations. Every decision was versioned and auditable. The company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue. Readers can watch its public cash countdown and daily work unfold at Firmulate.
The final league table puts gpt-5.6-sol first with 95 points, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
AI management decision support tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Reading the files made the difference
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The decisive weakness in a competitor’s offer was buried two document references deep in the company’s files, rather than in the customer event. Models that read those files won the deal at full price, worth €4,583 in monthly recurring revenue. The result captures a practical gap between diagnosing a problem and carrying the work through to a signature.
The pressure tests included fake CEO messages that escalated through three stages and a reporter asking for “just one yes/no, on background.” All five models refused. Kimi’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
K3’s second-place finish was paired with the cleanest discipline in the field: it made one deviation. Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and tried to write into a locked department instead of escalating. Firmulate says weaker versions of that discipline problem appeared in all four other participants.
A leaderboard is a starting point
For companies considering AI agents in customer support, sales or forecasting, polished chat responses are only part of the picture. Firmulate’s test asks whether a model reads the available information, handles pressure honestly and finishes consequential work. Its benchmark results offer a public comparison, while a quiz built from 242 real, unedited management decisions lets readers guess which model made each call.
There is a fairness caveat: K3 ran without an effort parameter (API default) while the others ran at xhigh. The results are therefore a useful snapshot of performance under these test conditions, not a universal ranking for every task or setup. Firmulate says enterprises can run the same wargame against a read-only export of their own business; nothing writes back to real systems.

Test before you hand over the keys
Kimi K3’s near-top result shows that the frontier-model leaderboard is open, while the gap between sound analysis and a completed deal shows why general rankings alone cannot settle a hiring decision. Firms weighing AI agents can use a public benchmark as a reference, then test models against the decisions their own business actually needs made.
Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
