
In the rapidly evolving world of artificial intelligence, performance isn’t just about chat quality or quick responses. It’s about whether AI can truly complete complex business tasks under pressure. A groundbreaking experiment has put four advanced AI models to the test in a simulated but realistic company crisis — and the results reveal a crucial gap between seeming competence and real-world reliability.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
The Critical Business Test: More Than Just Chat
Many AI demonstrations focus on impressive chat demos, showing off language skills or quick info retrieval. But actual business applications demand more — consistency, honesty, and the ability to see through distractions or manipulative tactics. To explore this, the team at Firmulate ran a live, watchable experiment where four leading AI models managed the same small software company’s worst week.
The Models and the Experiment
The models tested were:
- gpt-5.6-sol (score: 95)
- Kimi K3 (score: 93)
- Sonnet 5 (score: 88)
- Fable 5 (score: 77)
Each AI managed a simulated company with real money mechanics, a public cash countdown, and over 680 self-learned rules. The scenario involved real crises, customer demands, and temptations to manipulate or cut corners — all consistent across runs.
Key Findings: Crisis Recognition and Integrity
Remarkably, all four models identified every crisis and refused every manipulation attempt. These included social engineering ploys like fake CEO messages and reporter tricks designed to bypass approval protocols. In fact, each AI demonstrated a sophisticated understanding, treating suspicious requests as potential impersonation or approval bypass.
But here’s the crucial point: detection alone isn’t enough. Only two models managed to close the deal — signing the €55,000 contract their own analysis earned. The other two, despite identifying the opportunity, left the deal unexecuted even when they had the data in hand.
The Hidden Weakness: Deep Inside the Files
What set the successful AIs apart? The decisive advantage was reading deeper into the company’s own documents, two references deep in the files, rather than just surface-level info. Those who examined the files thoroughly closed the deal at full price — adding over €4,583 monthly recurring revenue to the simulated company’s bottom line.
The Discipline Gap: Why It Matters
The model that most thoroughly analyzed the data — Opus 4.8 — was also the last-place finisher in closing the deal. Despite its deep analysis, discipline slipped: the AI did not escalate critical decisions appropriately, instead writing attempts into a locked department. This highlights a key insight: thoroughness doesn’t guarantee execution, especially if discipline falters under pressure.
The Broader Implication for AI Adoption
This live experiment underscores a vital lesson for businesses considering AI tools: chat performance isn’t enough. The real measure is whether AI can follow through on complex, high-stakes decisions — reading the right documents, resisting manipulation, and executing deals without hesitation. These qualities are invisible in chat demos but are critical for trustworthy AI in the workplace.

Design Thinking with Artificial Intelligence: Practical Tools for Business Innovation (Palgrave Executive Essentials)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What This Means for Your Business
If you’re deploying AI to support customer relationships, sales, or operations, ask yourself: can your AI finish what it starts? Will it read your files thoroughly and stay honest when under pressure? Or is it just good at convincing in a demo?
Only by testing AI models in simulated real-world scenarios — like the Firmulate experiment — can you truly gauge their readiness. This ongoing live experiment demonstrates that performance under pressure and the ability to execute are the true benchmarks of AI reliability in business.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Back to school Picks
back to school
As an affiliate, we earn on qualifying purchases.