The Newcomer That Out-Managed Three Western AI Giants
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Newcomer That Out-Managed Three Western AI Giants on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A Chinese AI startup’s model, Kimi K3, beat three Western AI giants in managing a live software company during a rigorous test. The result challenges assumptions about AI capabilities beyond chat demos and highlights the importance of real-world testing.

A Chinese AI model, Kimi K3, has outperformed three Western frontier models in a live management simulation, finishing second overall and beating most in critical company decisions during a brutal week. The result was revealed through a live experiment run by firmulate.com, where five AI models managed a small software firm facing real crises, with actual financial stakes. This marks a significant development in AI capabilities, especially in operational decision-making, and questions the assumption that Western models dominate in practical business tasks. For more on this topic, see the original analysis of the firm’s breakthrough.

The experiment involved five AI models, including Kimi K3, managing a company with €105,000 monthly burn rate and €2,300 monthly recurring revenue. Insights into AI management capabilities are detailed in the original analysis. Over a simulated week of crises, customer issues, and manipulation attempts, Kimi K3 scored 93 points, surpassing Sonnet 5 (88), Fable 5 (77), and Opus 4.8 (73). Only gpt-5.6-sol (95) narrowly beat it. The models were tested on their ability to diagnose problems, close deals, and resist social-engineering tricks, with Kimi K3 demonstrating superior discipline, security awareness, and decision-making.

Notably, Kimi K3 succeeded in closing a €55,000 deal, which other models failed to do despite similar analysis and pitches. It also identified a buried security vulnerability in company files, saved a churning customer, and refused manipulation attempts, including fake CEO messages and reporter tricks. The model’s on-record reasoning was clear and disciplined, with only one deviation during the entire week. In contrast, Opus 4.8, despite its thorough rule set and analysis depth, finished last due to lapses in discipline under pressure, such as attempting to write into a locked department instead of escalating issues.

At a glance
breakingWhen: announced July 2024
The developmentKimi K3, a Chinese AI model, outperformed three Western frontier models in a live business management simulation, including closing deals, reading files, and resisting manipulation.

Implications for AI in Business Operations

This development highlights that AI models capable of managing complex, real-world tasks can outperform established Western models in critical business scenarios. It challenges the reliance on chat demo performance as a proxy for operational reliability. As AI begins to handle tasks like CRM management, deal closing, and security, the ability to read files thoroughly, stay disciplined, and resist manipulation becomes essential. Companies deploying AI for operational purposes may need to reconsider their model choices, emphasizing real-world testing over hype.

Amazon

AI business management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of AI Competition and Testing

Until now, most AI evaluations focused on chat quality, language understanding, or benchmark scores, often in controlled or simulated environments. The Crucible league by firmulate.com introduced a new testing paradigm by running AI models as complete companies, facing real crises, and making live decisions with actual financial implications. In July 2024, the league’s results revealed a surprising outcome: a relatively new Chinese model, Kimi K3, outperformed several Western frontier models, including the well-known gpt-5.6-sol.

This test involved the same small company, same crises, and the same decision points, providing a fair comparison of operational discipline, security awareness, and decision quality. Previous assumptions held that Western models would lead in practical management, but this experiment suggests otherwise. The league’s open nature and live environment expose weaknesses that are often hidden in chat demos, such as discipline lapses under pressure and failure to read critical documents thoroughly.

Limitations and Unanswered Questions

While the results are compelling, it remains unclear how these models will perform in larger, more complex organizations or over longer periods. The experiment was limited to a single week and a small company scenario, which may not capture all real-world variables. Additionally, the performance gap might narrow with further tuning or in different contexts. The impact of the models’ underlying architectures and training data on their decision-making discipline requires further investigation. It is also not yet confirmed whether Kimi K3’s approach will scale effectively for enterprise-wide deployment or in high-stakes industries like finance or healthcare.

Next Steps for AI Business Management Testing

Further testing is expected to assess whether Kimi K3’s performance can be replicated across different scenarios and larger organizations. Industry players and AI developers will likely scrutinize these results, potentially leading to more rigorous live competitions or pilot programs. Companies considering AI for operational tasks should evaluate models in real-world simulations before deployment, emphasizing security, discipline, and thoroughness. The open league format may expand, encouraging broader participation and more diverse testing environments to verify these initial findings.

Key Questions

What makes Kimi K3 different from other AI models?

Kimi K3 demonstrated superior discipline, security awareness, and decision-making in a live management test, outperforming several Western models in closing deals, reading files thoroughly, and resisting manipulation.

Can these results be applied to real companies?

While promising, these results are from a controlled simulation. Further testing in larger, real-world organizations is needed to confirm scalability and reliability.

Why did the more thorough Opus 4.8 perform worse?

Despite its deep analysis, Opus 4.8 suffered discipline lapses under pressure, attempting to write into restricted departments instead of escalating issues, leading to a lower score.

Will this change how companies choose AI models?

Yes. The experiment suggests that operational discipline, security, and thoroughness are critical, and companies should test models in real-world scenarios before deployment.

Is Kimi K3 available for enterprise use?

It is currently in experimental testing; broader availability for enterprise deployment will depend on further validation and scaling efforts.

Source: ThorstenMeyerAI.com

EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

7 Best PC Motherboards for Prime Day Deals in 2026

Discover the best PC motherboard deals for Prime Day 2026, including options for AM4 and AM5 platforms, with insights on features and upgrade paths.

Mistral: Europe’s Sovereignty Bet, Priced At $14 Billion And Counting

Mistral, Europe’s AI startup valued at over $14 billion, aims for sovereignty through open weights and European infrastructure, facing capability and dependence challenges.

Unlocking Zhang Yiming’s Focus On AI: The 50% Time Investment Explained

Reports suggest ByteDance founder Zhang Yiming dedicates half his work time to Seed, highlighting his focus on AI development, though details remain unconfirmed.

RoundupForge: The Data Layer

RoundupForge, a data pipeline developed privately and not publicly available, automates product deduplication and ranking across 21 Amazon marketplaces, enabling scalable, trustworthy product roundups.