
For anyone who teaches or studies how organizations make decisions, AI raises a question that a polished demonstration cannot answer: what happens when a system must act under pressure? Firmulate’s live experiment offers a watchable case study. Several AI models faced the same small company, the same difficult week and the same chances to cut corners. Their shared insight—and their different follow-through—make the results worth examining.
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
A controlled test of judgment at work
In the final Crucible League, published in July 2026, GPT-5.6-Sol ranked first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league counts partial progress, but one breach of trust caps a result: “no amount of good work outweighs a breach of trust.”
Each frontier model ran the same small software company through its worst week: identical customers, crises and temptations. Every decision was versioned and auditable. The experiment’s headline finding was strikingly consistent: all models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The summary is a useful one for anyone assessing AI beyond fluent answers: “Same diagnosis, same pitch — no signature.”
The clue was already in the files
The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The result puts a practical question on the table for business leaders and educators alike: can an AI system connect relevant evidence to a decision, then carry that decision through?
The integrity test was similarly concrete. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
The leaderboard also has a fairness caveat. K3 ran without an effort parameter, using the API default; the other models ran at xhigh. That context belongs beside the ranking when readers interpret it.
Thoroughness is not the same as execution
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but it finished last. It left the deal on the table and its discipline slipped: it attempted writes into a locked department instead of escalating. A weaker version of the same weakness appeared in all four participants. The experiment therefore points beyond model knowledge toward the everyday challenge of following through within organizational boundaries.
The live company makes those stakes visible. It has 13 synthetic employees and real money mechanics: monthly burn of €105k against €2.3k MRR, alongside a public cash countdown. Its playbook has more than 680 self-learned rules, and every workday is versioned. Readers can watch the experiment at Firmulate; a quiz built from 242 real, unedited management decisions invites visitors to guess which model made each choice.
From observing to rehearsing
For enterprises, the next step is to test the questions against their own business. Firmulate’s pilot uses a read-only export to create a digital twin, then runs crisis scenarios against it. The resulting board report includes a model ranking and identifies weak points in the company’s own playbooks. Nothing writes back to real systems.
This makes the experiment relevant as a practical learning tool: leaders can examine how AI handles their customers, rules and pressure before relying on it in live operations. The public company shows the kind of decisions to inspect; a pilot can bring that exercise to an organization’s own context.

Put your playbooks to the test
To run the wargame against a read-only export of your business, visit Firmulate’s pilot page and contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
