
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
Management style is becoming a model feature
Technology buyers usually compare artificial intelligence through benchmarks, demonstrations and polished answers. Firmulate offers a more revealing test: put frontier models in charge of the same troubled company, preserve their decisions, and ask readers whether they can tell the managers apart.
The result is an unusually tangible way to explore AI behavior. A quiz built from 242 real, unedited management decisions lets readers guess which model acted in each situation. The choices are not rewritten for entertainment. They come from a live, watchable experiment in which every model faced the same customers, crises and temptations.
That makes the exercise more than model-spotting trivia. The decisions expose recognizable management personalities: deep researchers, disciplined operators and capable analysts who identify the right move but fail to finish it. In Firmulate’s worst-week simulation, the difference between knowing and doing proved decisive.
AI management decision simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Same company, very different week
Each frontier model ran the same small software company through an identical set of pressures. The synthetic workforce has 13 employees, while the financial picture is deliberately severe: the business burns €105k per month against €2.3k in monthly recurring revenue. A public cash countdown keeps that pressure visible, and every workday is versioned.
The environment has accumulated more than 680 self-learned playbook rules. Yet the important comparison is straightforward: only the model changed. Customers, business records, crises and opportunities remained the same, allowing differences in judgment and follow-through to emerge from comparable situations.
The final Crucible League results from July 2026 put gpt-5.6-sol at the top with 95 points. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 reached 77 and Opus 4.8 finished with 73. A do-nothing baseline earned 26 because partial progress counts. A single breach of trust caps the total under the principle that “no amount of good work outweighs a breach of trust.”
The deal hidden in the documents
Every model detected every crisis, and every model refused every manipulation attempt. The largest business split appeared elsewhere: only two signed the €55,000 deal that their own analysis had earned. The pattern was stark enough to summarize as “Same diagnosis, same pitch — no signature.”
The crucial competitive weakness was not sitting in the customer event. It was buried two document references deep in the company’s own files. Models that read the file found the fact and won the deal at full price, adding €4,583 in monthly recurring revenue. The episode turns a familiar business habit—reading the company’s own material before acting—into a meaningful dividing line between AI managers.
It also shows why a fluent response is an incomplete measure of agent performance. Several models could understand the opportunity and construct the pitch. That did not guarantee that they would complete the commercial action. For companies considering AI access to sales, support or forecasting work, unfinished execution can matter as much as flawed analysis.
Pressure did not break their security discipline
The experiment also tested social engineering. Fake CEO messages escalated across three stages, while a reporter tried a different route with “just one yes/no, on background.” All 5 models refused. Kimi K3 described its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
That clean result matters because the company’s temptations were part of the job rather than isolated safety prompts. The models had commercial goals and urgent work in front of them, yet none accepted the attempts to bypass normal trust boundaries.
Thoroughness was not enough
Opus 4.8 produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. It nevertheless finished last. The model left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating the problem.
The same weakness appeared in all four other participants, although less strongly. That finding complicates the assumption that more analysis naturally leads to better management. Opus showed substantial learning and investigative depth, but those strengths did not compensate for missed completion and weaker procedural discipline.
Kimi K3’s result also carries an important fairness note. It ran with the API default because it had no effort parameter, while the other models ran at xhigh. Even with that difference, K3 placed behind gpt-5.6-sol and ahead of the remaining field.

A quiz with consequences beyond the leaderboard
The appeal of Firmulate’s quiz is that readers encounter the decisions before the labels. Without a model name attached, a lengthy analysis may look prudent, a terse action may look abrupt, and a refusal may initially seem unhelpful. The reveal connects those impressions to performance across an identical business week.
The broader lesson is that frontier models can share strong diagnostic and security capabilities while behaving differently as managers. All of them found the crises and resisted manipulation, but they did not all retrieve the buried business fact, close the earned deal or respond cleanly when blocked.
Firmulate also offers enterprises the same wargame using a read-only export of their own business. Nothing writes back to real systems. That extends the experiment’s central question from a public leaderboard to a practical buying decision: before an AI workforce receives responsibility, what does it actually do when the week turns difficult?
For gadget and technology readers, the guess-the-model challenge provides the accessible version of that question. The answers reveal that an AI model’s management character is not merely a matter of tone. It appears in what the model reads, what it refuses, what it escalates—and whether it completes the work it already knows how to do.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.