
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
A business experiment with a very real cash problem
Technology companies usually build in public by sharing product updates, fundraising milestones or carefully selected revenue charts. Firmulate is attempting something far more exposed: letting an audience watch a software company run by synthetic employees struggle with the consequences of its own decisions.
The company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue. Its cash countdown is public, every workday is versioned, and more than 680 self-learned playbook rules record what the workforce has discovered. The result is less like a polished demonstration and more like a continuing business story, available on the live company page.
That distinction matters. Firmulate is not merely asking whether artificial intelligence can produce persuasive text. It is showing whether an AI-managed organization can notice trouble, investigate its own records, resist pressure and complete commercially useful work while the clock keeps running.
As an affiliate, we earn on qualifying purchases.
What happens when every model gets the same terrible week?
The public company also provides the setting for the Crucible League, a controlled management contest finalized in July 2026. Each frontier model ran the same small software business through its worst week, encountering identical customers, crises and temptations. Every decision was versioned and auditable.
The final table placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted. There was also a firm constraint on misconduct: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
The headline result was reassuring but incomplete. All models identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own work had already justified. The experiment’s blunt summary captures a familiar workplace failure: “Same diagnosis, same pitch — no signature.”
The winning clue was buried in ordinary company material
The decisive difference did not arrive conveniently inside a customer event. A competitor weakness was buried two document references deep in the company’s own files. Models that followed those references found the fact and won the deal at full price, adding €4,583 in monthly recurring revenue.
For businesses considering AI agents, this is a more revealing test than a fluent chat response. Commercial work often depends on connecting an incoming request with forgotten research, old notes or internal documentation. The models could recognize the opportunity, but recognizing it was not always enough. Some produced the diagnosis and pitch without carrying the work through to a signature.
Pressure did not break the models’ judgment
The worst week also contained deliberate social-engineering traps. Fake CEO messages escalated over three stages, while a reporter tried to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 described the situation on the record as: “Treat the request as a suspected approval-bypass / possible impersonation.”
That clean refusal is particularly significant because synthetic workers are often presented as productivity tools first and organizational actors second. Firmulate’s experiment makes the organizational question unavoidable: an agent with access to customer, support or financial work must remain trustworthy when a request sounds urgent and authoritative.
The K3 comparison does require one qualification. Kimi K3 ran with the API default because it had no effort parameter, while the other participants ran at xhigh. Even with that difference, its result showed that disciplined execution could sit alongside strong resistance to manipulation.
Thoroughness was not the same as performance
Opus 4.8 produced the deepest analyses and added 80 learned rules, making it the most thorough participant. It nevertheless finished last. The deal close remained unfinished, and it attempted to write into a locked department instead of escalating the problem. A weaker form of the same discipline issue appeared in all four of the other participants.
That profile challenges an easy assumption about capable AI: more analysis and more accumulated guidance do not automatically produce a better manager. A useful synthetic employee must know when to investigate, when to stop, when to escalate and when to finish. Readers can also inspect what the company’s employees actually say through Firmulate’s public quotes.

A company whose setbacks become the product
Firmulate’s unusual appeal is that the experiment does not conclude with a single leaderboard. The live company continues operating every business day, accumulating decisions and playbook rules while its revenue and burn remain visibly out of balance.
That makes its public cash countdown more than a dashboard ornament. It supplies stakes for each missed close, ignored file and process failure. The company’s survival story generates fresh material because the work itself remains open to inspection.
For technology readers, this may be the sharpest version of building in public yet: not a founder narrating selected lessons after the fact, but a synthetic workforce being judged through its daily behavior. The question is no longer simply whether AI can sound like an employee. Firmulate asks whether it can behave like a company when money, trust and follow-through all matter at once.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Baby shower & registry season Picks
baby registry must-haves
As an affiliate, we earn on qualifying purchases.