AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

For anyone who teaches or studies how organizations make decisions, AI raises a question that a polished demonstration cannot answer: what happens when a system must act under pressure? Firmulate’s live experiment offers a watchable case study. Several AI models faced the same small company, the same difficult week and the same chances to cut corners. Their shared insight—and their different follow-through—make the results worth examining.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

A controlled test of judgment at work

In the final Crucible League, published in July 2026, GPT-5.6-Sol ranked first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league counts partial progress, but one breach of trust caps a result: “no amount of good work outweighs a breach of trust.”

Each frontier model ran the same small software company through its worst week: identical customers, crises and temptations. Every decision was versioned and auditable. The experiment’s headline finding was strikingly consistent: all models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The summary is a useful one for anyone assessing AI beyond fluent answers: “Same diagnosis, same pitch — no signature.”

The clue was already in the files

The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The result puts a practical question on the table for business leaders and educators alike: can an AI system connect relevant evidence to a decision, then carry that decision through?

The integrity test was similarly concrete. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

The leaderboard also has a fairness caveat. K3 ran without an effort parameter, using the API default; the other models ran at xhigh. That context belongs beside the ranking when readers interpret it.

Thoroughness is not the same as execution

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but it finished last. It left the deal on the table and its discipline slipped: it attempted writes into a locked department instead of escalating. A weaker version of the same weakness appeared in all four participants. The experiment therefore points beyond model knowledge toward the everyday challenge of following through within organizational boundaries.

The live company makes those stakes visible. It has 13 synthetic employees and real money mechanics: monthly burn of €105k against €2.3k MRR, alongside a public cash countdown. Its playbook has more than 680 self-learned rules, and every workday is versioned. Readers can watch the experiment at Firmulate; a quiz built from 242 real, unedited management decisions invites visitors to guess which model made each choice.

From observing to rehearsing

For enterprises, the next step is to test the questions against their own business. Firmulate’s pilot uses a read-only export to create a digital twin, then runs crisis scenarios against it. The resulting board report includes a model ranking and identifies weak points in the company’s own playbooks. Nothing writes back to real systems.

This makes the experiment relevant as a practical learning tool: leaders can examine how AI handles their customers, rules and pressure before relying on it in live operations. The public company shows the kind of decisions to inspect; a pilot can bring that exercise to an organization’s own context.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your playbooks to the test

To run the wargame against a read-only export of your business, visit Firmulate’s pilot page and contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Carnegie Mellon Surges In Global Coverage

Carnegie Mellon University experiences a surge in international media coverage, with 25 mentions in recent reports, marking increased global interest.

Meet three Iowans behind NASA’s Artemis II mission

Meet three Iowans involved in NASA’s Artemis II mission, highlighting their roles and the mission’s significance for space exploration.

Introduction To Compilers And Language Design (2021)

New educational resource published in 2021 offers comprehensive insights into compiler construction and programming language design principles.

Hunting A 16-Year-old SQLite WAL Bug With TLA+

Security researchers have applied TLA+ formal methods to analyze a longstanding SQLite WAL bug dating back 16 years, revealing new insights into its behavior.