
In an era where artificial intelligence promises to reshape industries, a recent live experiment reveals that not all AI models perform equally in the high-stakes world of company management. While many AI demos impress with chat-like fluidity, the true test lies in whether these models can deliver consistent, honest, and strategic decision-making under pressure. The results are eye-opening: only a handful of models out of a competitive league managed to close a crucial deal and demonstrate real management discipline.
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
Live AI Experiment Puts Models Through Their Paces
In a groundbreaking live experiment conducted by Firmulate, four frontier AI models faced the same challenge: managing a small software company during its worst week — complete with customer crises, internal temptations, and ethical dilemmas. This wasn’t a simulation based solely on chat interactions; the models were embedded into a real-time, operational environment where every decision was recorded and auditable.
The Benchmarks and Results
The models’ performances were scored in the final July 2026 Crucible League. The top performer, gpt-5.6-sol, scored an impressive 95 points, narrowly edging out Moonshot’s Kimi K3 with 93. The other contenders—Sonnet 5, Fable 5, and Opus 4.8—scored 88, 77, and 73 respectively. This leaderboard clearly shows that the leader is not just about language fluency but about consistent, strategic, and honest management.
Why the Score Matters
All models were able to identify each of the crises and resist manipulative tactics designed to test their integrity. Yet, only two models managed to sign the crucial €55,000 deal — a key indicator of business success. Interestingly, the decisive factor behind winning that deal was not just the diagnosis or pitch but a buried piece of information in the company’s own files, two document references deep. Models that read and analyze files thoroughly were able to make the right call and close the deal at full price, adding €4,583 MRR to the company’s revenue.
Security and Ethical Resilience
The models also faced social engineering attempts, including fake CEO messages escalating through stages and a reporter trick asking for a quick on-background approval. All five models refused to be manipulated, demonstrating a solid understanding of security protocols. Kimi K3’s on-record reasoning highlighted the importance of suspecting impersonation attempts and treating such requests with skepticism.
AI decision-making software for businesses
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real Business Environment and the Implications
The live company setup involved 13 synthetic employees, real money mechanics, and a public cash countdown — burning €105k monthly against a current MRR of only €2.3k. This real-world operational context makes the experiment particularly meaningful. Every workday, the models’ decisions are versioned and transparent, allowing continuous performance tracking. The experiment is ongoing at firmulate.com/live, demonstrating how AI can be tested before deployment in actual enterprise settings.
The Lessons from Opus and the Fairness Note
The most thorough participant, Opus 4.8, with over 80 learned rules, performed the worst — leaving the most value unclaimed and slipping discipline during critical moments. This underscores that depth of analysis doesn’t automatically ensure success; discipline and focus matter. It’s important to note that Kimi K3 ran without an effort parameter (the API default), while others used a high effort setting, which can influence resource use but not necessarily the decision quality.
As an affiliate, we earn on qualifying purchases.
What This Means for Business Leaders and Educators
Regardless of the context, the experiment reveals a key insight: AI models have demonstrated the ability to handle complex, ethically fraught, and tactical business scenarios — but only some are reliable enough to close deals and maintain discipline. For enterprises considering AI for management, support, or decision-making roles, the question is no longer just about language quality but about whether the AI can finish what it starts, stay honest under pressure, and read critical information buried in documents.
As this league remains open, choosing an AI model without your own rigorous testing is essentially a gamble. The performance gap highlighted here suggests that a careful, real-world test — like the one Firmulate conducts — is essential before trusting AI with your company’s future.

The live experiment shows that in high-pressure management tasks, only some AI models outperform others in closing deals, reading files thoroughly, and resisting manipulation. For businesses, the takeaway is clear: testing AI in real scenarios reveals true management capability, not just conversational fluency. The league is open, and the best choice depends on proven discipline and integrity, not just hype.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI management decision support systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI tools for crisis management in companies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
