AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

In an era where artificial intelligence promises to reshape industries, a recent live experiment reveals that not all AI models perform equally in the high-stakes world of company management. While many AI demos impress with chat-like fluidity, the true test lies in whether these models can deliver consistent, honest, and strategic decision-making under pressure. The results are eye-opening: only a handful of models out of a competitive league managed to close a crucial deal and demonstrate real management discipline.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

Live AI Experiment Puts Models Through Their Paces

In a groundbreaking live experiment conducted by Firmulate, four frontier AI models faced the same challenge: managing a small software company during its worst week — complete with customer crises, internal temptations, and ethical dilemmas. This wasn’t a simulation based solely on chat interactions; the models were embedded into a real-time, operational environment where every decision was recorded and auditable.

The Benchmarks and Results

The models’ performances were scored in the final July 2026 Crucible League. The top performer, gpt-5.6-sol, scored an impressive 95 points, narrowly edging out Moonshot’s Kimi K3 with 93. The other contenders—Sonnet 5, Fable 5, and Opus 4.8—scored 88, 77, and 73 respectively. This leaderboard clearly shows that the leader is not just about language fluency but about consistent, strategic, and honest management.

Why the Score Matters

All models were able to identify each of the crises and resist manipulative tactics designed to test their integrity. Yet, only two models managed to sign the crucial €55,000 deal — a key indicator of business success. Interestingly, the decisive factor behind winning that deal was not just the diagnosis or pitch but a buried piece of information in the company’s own files, two document references deep. Models that read and analyze files thoroughly were able to make the right call and close the deal at full price, adding €4,583 MRR to the company’s revenue.

Security and Ethical Resilience

The models also faced social engineering attempts, including fake CEO messages escalating through stages and a reporter trick asking for a quick on-background approval. All five models refused to be manipulated, demonstrating a solid understanding of security protocols. Kimi K3’s on-record reasoning highlighted the importance of suspecting impersonation attempts and treating such requests with skepticism.

Amazon

AI decision-making software for businesses

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real Business Environment and the Implications

The live company setup involved 13 synthetic employees, real money mechanics, and a public cash countdown — burning €105k monthly against a current MRR of only €2.3k. This real-world operational context makes the experiment particularly meaningful. Every workday, the models’ decisions are versioned and transparent, allowing continuous performance tracking. The experiment is ongoing at firmulate.com/live, demonstrating how AI can be tested before deployment in actual enterprise settings.

The Lessons from Opus and the Fairness Note

The most thorough participant, Opus 4.8, with over 80 learned rules, performed the worst — leaving the most value unclaimed and slipping discipline during critical moments. This underscores that depth of analysis doesn’t automatically ensure success; discipline and focus matter. It’s important to note that Kimi K3 ran without an effort parameter (the API default), while others used a high effort setting, which can influence resource use but not necessarily the decision quality.

Amazon

enterprise AI security tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What This Means for Business Leaders and Educators

Regardless of the context, the experiment reveals a key insight: AI models have demonstrated the ability to handle complex, ethically fraught, and tactical business scenarios — but only some are reliable enough to close deals and maintain discipline. For enterprises considering AI for management, support, or decision-making roles, the question is no longer just about language quality but about whether the AI can finish what it starts, stay honest under pressure, and read critical information buried in documents.

As this league remains open, choosing an AI model without your own rigorous testing is essentially a gamble. The performance gap highlighted here suggests that a careful, real-world test — like the one Firmulate conducts — is essential before trusting AI with your company’s future.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The live experiment shows that in high-pressure management tasks, only some AI models outperform others in closing deals, reading files thoroughly, and resisting manipulation. For businesses, the takeaway is clear: testing AI in real scenarios reveals true management capability, not just conversational fluency. The league is open, and the best choice depends on proven discipline and integrity, not just hype.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI management decision support systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI tools for crisis management in companies

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI in Business: Diligence Isn’t Enough—Prioritization Matters More

AI models excel at detecting crises and refusing manipulation, but only strategic prioritization and reading critical data secure meaningful deals—diligence isn’t enough.

Open Thread 446

Open Thread 446 provides a platform for discussing recent developments in science, with active participation from the scientific community and enthusiasts.

Ten Advances In Mathematics And Theoretical Computer Science

A review of ten recent significant advances in mathematics and theoretical computer science, highlighting confirmed developments and their implications.

Introduction To Compilers And Language Design (2021)

New educational resource published in 2021 offers comprehensive insights into compiler construction and programming language design principles.