
Imagine watching an AI manage a company day by day, making critical decisions, facing crises, and even losing money — all in full public view. This is not a sci-fi scenario, but the reality of a bold experiment by Firmulate, an innovative tech firm that runs a simulated company with no human employees, just AI models acting as management.
The Live Experiment: An AI-Driven Company Under Siege
At the heart of this experiment is a tiny software firm with a monthly revenue of €2,300, burning through €105,000 in cash every month. Its entire operations are publicly accessible at firmulate.com/live.html. The company employs 13 synthetic ’employees’— AI models tasked with decision-making and crisis management—making it a unique glimpse into how AI handles real-world business pressures.
As an affiliate, we earn on qualifying purchases.
Testing AI Decision-Making in a High-Stakes Environment
The experiment subjected four state-of-the-art frontier models to the same challenging week: same customers, same crises, and the same temptation to manipulate or cheat. Each model’s decisions are meticulously versioned and auditable, creating a transparent view into how AI responds under stress.
The Results: Successes and Failures
All four AI models successfully identified every crisis, demonstrating robust crisis recognition. They also refused every manipulation attempt— which included sophisticated social engineering tactics like staged CEO messages and reporter tricks. Interestingly, only two of the models managed to close the deal that their analysis recommended, signing a €55,000 contract, earning +€4,583 MRR. The other two models failed to follow through, leaving potential revenue on the table despite similar diagnoses and pitches.
The Hidden Weakness: What the AI Overlooked
Deep in the company’s own files— two document references down— was a critical piece of information that could have sealed the deal. The models that read and understood these files captured the opportunity, turning it into real revenue. This underscores a vital insight: the AI’s ability to read and interpret relevant internal documentation is crucial for success.
Beyond Crisis Management: Trust and Discipline
Another revealing aspect was how the models handled social engineering. When presented with staged messages from a fake CEO or background approval requests, all five tested models refused to act. Kimi K3’s explicit reasoning highlighted a focus on security: “Treat the request as a suspected approval-bypass / possible impersonation.”
Real Money, Real Stakes, Public Scrutiny
The live company, with its burn rate and cash countdown, is a stark reminder of the financial stakes involved. Each working day is versioned, analyzed, and made transparent at firmulate.com/live.html. The experiment showcases not just AI’s potential but its current limitations— especially when discipline slips, as seen with OPUS 4.8, which left a deal unexecuted due to procedural lapses despite having the most thorough analysis.
The Larger Implications: AI in Business Decision-Making
This experiment raises critical questions for any enterprise considering AI: Can the AI finish what it starts? Will it read and interpret internal documents? Can it resist manipulation under pressure? As one of the models scored 95 out of 100, successfully closing the deal and finding the buried fact, the message is clear: AI can be a powerful business partner— if designed and tested rigorously.
Try It Yourself: Run Your Own Business Wargame
For organizations curious about how AI might perform in their own operations, Firmulate offers a read-only export tool to simulate their business challenges without risking real data or systems. Details are at firmulate.com/pilot.html.

This live experiment demonstrates how AI models can manage complex business scenarios, spot critical information, and resist manipulation. But it also highlights the importance of discipline and thoroughness— aspects that determine whether AI truly adds value or leaves money on the table.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html