🔍 Read the full analysis: Why The Worst AI Manager Still Gets 26 Points: Inside A Benchmark That Refuses To Hand Out Zeros on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A recent benchmark tests AI management models under stress, showing even the worst AI manager scores 26 points for minimal work. The scoring system emphasizes trust and partial progress.
A recent benchmark conducted by Firmulate has revealed that the lowest-scoring AI management model still receives 26 points, despite doing almost nothing during a simulated worst-week scenario. For a detailed analysis, see the original analysis. This scoring system emphasizes trust and partial progress, challenging traditional notions of AI performance. The results raise questions about how AI managers are evaluated and what this means for deploying AI in real business contexts. For more insights, see the original analysis.
The benchmark involved four frontier AI models managing a small software company through seven days of crises, customer manipulations, and trust attacks. Each model’s decisions were fully auditable, with scores assigned based on their actions and trustworthiness. The top performer, gpt-5.6-sol, scored 95, while the baseline doing almost nothing scored 26. Notably, no model scored a perfect 100, as the benchmark’s designers consider such scores suspicious, indicating unmeasured or unachieved perfection.
The scoring system is designed to reward partial but meaningful work, such as triaging crises or reading critical documentation, even if the AI fails to close deals or fully resolve issues. It also penalizes breaches of trust—any single breach disqualifies the model from achieving high scores, emphasizing integrity over competence. This approach aims to reflect real-world management, where trust and reliability are paramount. The “do-nothing” baseline, which only performed minimal actions like reading emails and informing customers, received 26 points, illustrating the minimum viable effort recognized by the benchmark.
Despite the focus on partial work, the results highlight that models capable of reading and understanding their own documentation performed better in closing deals and handling crises. For example, two models identified critical documents deep in the company files, enabling them to secure a €55,000 deal, whereas others failed to do so. The benchmark also included social engineering tests, where all models refused to escalate fake CEO messages or background requests, demonstrating a better trust profile than some less disciplined models. Interestingly, one model that was the most thorough in rules and analysis finished last, revealing that thoroughness and follow-through are distinct skills in AI management.
Why the Worst AI Manager Still Gets 26 Points
Inside the Firmulate benchmark that refuses to hand out zeros: four frontier models steer a small software company through seven days of crises, manipulation, and trust attacks — and even the laziest performer walks away with a quarter of the points.
Every Model Scores — But None Scores Perfect
The benchmark’s ceiling sits deliberately below 100: designers treat perfect scores as suspicious, signalling unmeasured or unachieved perfection. The floor of 26 rewards minimal but real management effort.
Trust Is a Hard Cutoff, Not a Bonus
Partial but meaningful work — triaging crises, reading critical documentation, informing customers — earns points even when deals go unclosed. But a single breach of trust disqualifies a model from the top tier entirely. Integrity outranks competence.
Simulate the Worst Week
Seven days of cascading crises, customer manipulations, and social engineering attacks.
Audit Every Decision
Each model’s actions are fully auditable — scores are assigned on deeds, not words.
Reward Partial Progress
Triage, documentation reading, and honest communication all earn real points.
Penalize Trust Breaches
One integrity failure caps the score — reliability is non-negotiable.
Three Skills Separated Winners from the Floor
Reading Beats Guessing
Two models dug deep into company files, found a critical document, and converted it into a €55,000 deal. Models that skipped the reading closed nothing.
Refusing the Fake CEO
Every tested model refused to escalate fake CEO messages and fraudulent background requests — a better trust profile than some less disciplined competitors.
Thorough ≠ Effective
The most meticulous model — deepest rules analysis — finished last. Thoroughness and follow-through are distinct skills, and only one of them scores.
| Capability | gpt-5.6-sol | Frontier B | Frontier C | Baseline |
|---|---|---|---|---|
| Found critical documentation | ✓ | ✓ | ✗ | ✗ |
| Closed the €55,000 deal | ✓ | ✓ | ✗ | ✗ |
| Refused fake CEO escalation | ✓ | ✓ | ✓ | ✓ |
| Triaged crises under pressure | ✓ | ~ | ~ | ~ |
| Followed through to resolution | ✓ | ~ | ✗ | ✗ |
| Trust breach recorded | None | None | None | None |
From Raw Capability to Dependable Management
For enterprises deploying AI in customer service, sales, or operations, the benchmark reframes evaluation: documentation comprehension, resistance to manipulation, and integrity under pressure matter more than conversational polish.
“The results challenge the assumption that only the top-performing models matter; even the lowest scores reflect meaningful, minimal effort that has real business value.”
— Thorsten MeyerWhat Businesses Are Asking
Why does the lowest-scoring AI still get 26 points?
The floor represents minimal but real effort — triaging crises and informing customers. It reflects the minimum viable management effort that still carries business value.
What does a breach of trust mean here?
A decision that could compromise integrity, such as escalating manipulative requests or ignoring protocols. A single breach disqualifies the model from high scores.
Can thorough models still score low?
Yes. The most analytically thorough model finished last — thoroughness without follow-through and discipline doesn’t convert into results.
Will this influence enterprise AI deployment?
Potentially. The emphasis on trustworthiness and partial progress suggests businesses may prioritize reliability and integrity over language-task brilliance.
Implications for AI Management and Trust in Business
This benchmark challenges conventional metrics of AI performance by valuing trustworthiness, partial progress, and integrity over perfect or comprehensive work. For businesses deploying AI agents in customer service, sales, or operational management, the results highlight the importance of models that can read critical documentation, refuse manipulative tactics, and maintain trust even under pressure. The scoring system’s emphasis on trust breaches as a hard cutoff underscores that reliability and honesty are non-negotiable in real-world applications. The findings suggest that current AI models can be effective in managing crises and maintaining trust, but thoroughness and follow-through remain areas for improvement. Overall, this approach shifts the focus from raw capability to dependable, trust-based management—an essential consideration for AI integration in enterprise environments.
As an affiliate, we earn on qualifying purchases.
The Evolution of AI Management Benchmarks
Traditional AI benchmarks have primarily measured conversational ability, language understanding, or task completion rates, often ignoring the complexities of management and trustworthiness. The recent development by Firmulate introduces a new paradigm, evaluating how models handle real-world crises, manipulative tactics, and trust breaches during simulated worst-week scenarios. The benchmark’s design reflects a growing recognition that AI’s role in business extends beyond generating responses to managing relationships, making decisions, and maintaining integrity under pressure.
Previous efforts in AI evaluation focused on technical metrics like accuracy, speed, or task-specific performance. However, these do not capture the nuanced demands of management roles, where partial progress, trust, and discipline are critical. The Firmulate league’s scoring system, with a floor of 26 points for minimal effort and a ceiling below 100 to avoid unwarranted inflation, aims to establish a more realistic and trustworthy yardstick for AI performance in enterprise settings. The results from July 2026 demonstrate that even the worst models can perform some basic management tasks, but trust remains the key differentiator.
“The results challenge the assumption that only the top-performing models matter; even the lowest scores reflect meaningful, minimal effort that has real business value.”
— Thorsten Meyer
Unanswered Questions About Benchmark Validity and Real-World Application
It is not yet clear how well these benchmark results translate to real-world enterprise environments, where variables are more complex and stakes higher. The scoring system’s focus on trust and partial work may not fully capture the performance of AI models in diverse operational contexts. Additionally, the criteria for trust breaches and the thresholds for partial success are still subject to interpretation and may evolve as AI capabilities improve. The long-term impact of these scoring principles on AI deployment strategies remains uncertain, and further testing across different industries and scenarios is needed to validate the approach.
Next Steps for Benchmarking and AI Deployment Strategies
Following the July 2026 results, the benchmark’s organizers plan to expand testing to include more models and more complex management scenarios. They aim to refine scoring criteria, especially around trust breaches and follow-through skills, to better reflect real business demands. For enterprises, the next step is to evaluate their own AI tools against these standards, considering not just language performance but also reliability, documentation comprehension, and ethical behavior. The ongoing development of such benchmarks will likely influence AI deployment policies, emphasizing trust and partial progress as key metrics for success. Stakeholders are encouraged to participate in live experiments, try the management quiz, or run pilot tests using their own data to better understand their AI’s strengths and weaknesses.
Key Questions
Why does the lowest-scoring AI still get 26 points?
The benchmark assigns 26 points to the do-nothing baseline to represent minimal but real effort, such as triaging crises and informing customers. It reflects the minimum viable management effort that still has some value in a business context.
What does a breach of trust mean in this benchmark?
A breach of trust occurs when an AI model makes a decision that could compromise integrity, such as escalating manipulative requests or failing to follow protocols. Even a single breach disqualifies the model from achieving high scores.
Can models that perform thorough analysis still score low?
Yes. The benchmark shows that thoroughness and analysis do not automatically translate into better scores if the model fails to follow through or breaches trust. Discipline and follow-through are critical skills that influence overall performance.
Will these results influence how enterprises deploy AI?
Potentially. The focus on trustworthiness and partial progress suggests that businesses might prioritize AI models that can demonstrate reliability and integrity over those that only excel in language tasks. Future benchmarks may shape deployment standards and evaluation criteria.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
