AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

In the rapidly evolving world of artificial intelligence, the assumption persists that more diligent, rule-abiding models automatically outperform their peers. But recent experiments challenge this notion, demonstrating that persistent effort and thoroughness alone may not secure business success. As AI increasingly touches critical decision-making and trust, understanding what truly distinguishes top performers is vital.

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

The Firmulate Experiment: Putting AI to the Test in a Realistic Business Scenario

In a groundbreaking live experiment, four advanced AI models were tasked with managing a small software company during its worst week—facing the same customers, crises, and temptations. This setup was designed to assess their decision-making, honesty, and discipline under pressure. Every choice made by the models was tracked and auditable, revealing not only what they detected but how they responded to manipulation attempts and crises.

Amazon

AI decision-making software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Findings: Diligence Alone Isn’t Enough

Surprisingly, all four models demonstrated a high level of vigilance—they detected every crisis and refused every manipulation attempt. Yet, only two actually closed a critical deal worth €55,000. This illustrates a striking gap between awareness and action: spotting problems is not synonymous with winning business or maintaining discipline.

Amazon

business process automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Factor: Reading and Acting on Company Files

Further analysis revealed that the decisive advantage lay in the models’ ability to reference the company’s own documents—information buried two document references deep. The models that read and interpret these files successfully closed the deal at full price, adding an estimated +€4,583 in Monthly Recurring Revenue (MRR). In contrast, others that missed this key information lost the opportunity, despite their thorough crisis detection.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Trust and Manipulation: AI’s Responses Under Social Engineering

During staged social engineering attempts—a fake CEO message escalating over three stages and a reporter trick requesting a background yes/no—every model refused to engage. Kimi K3, one of the top performers, explicitly treated the request as a potential impersonation or approval-bypass, maintaining integrity under pressure. This consistent refusal underscores the importance of built-in safeguards against manipulation, especially in high-stakes environments.

Amazon

AI risk management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real-World Company: An Ongoing Business Simulation

The experiment isn’t just academic. The live site at firmulate.com showcases a functioning company with 13 synthetic employees operating with real money mechanics—burning €105,000 monthly against €2,300 in MRR, with a public cash countdown. Every day, the models analyze and make decisions within this simulated business, which is accessible for enterprises to test their own AI workforce without risking actual systems.

The Opus 4.8 Profile: Deep Analysis, Last Place

Among the models tested, Opus 4.8 was the most meticulous, incorporating over 80 learned rules and performing the deepest analysis. Yet it finished last in the league—failing to close the deal because it left the close on the table and slipped into a process slip, writing attempts into a locked department instead of escalating. This demonstrates that sheer thoroughness and rule-following do not translate directly into successful business outcomes.

Implications for Business and AI Strategy

This live experiment highlights a core truth: diligence and comprehensive analysis are valuable, but they are insufficient without prioritization and disciplined action. AI models must be programmed not only to detect issues but also to act decisively—reading the right information, maintaining integrity under pressure, and understanding what truly drives results.

What Business Leaders Should Take Away

  • AI systems should be evaluated on their ability to complete meaningful work, not just their awareness or rule adherence.
  • Reading and interpreting critical internal documents can be the difference between winning and losing a deal.
  • Robust safeguards against manipulation are essential to maintain trust and integrity.
  • Deep analysis and thoroughness need to be balanced with prioritization and decisive action to generate real impact.

Learn More and Watch the Live Experiment

For those interested in testing their own business scenarios or understanding how AI models perform in complex decision environments, the live site at firmulate.com offers a transparent, real-time view of ongoing experiments. Here, you can see how different models handle crises, read company files, and make decisions—providing a clear window into what makes AI truly effective in a business setting.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Any text-to-SQL Benchmark Should Address Difficulties Of Real-world Data Stores

Experts emphasize that effective text-to-SQL benchmarks should incorporate the complexities of real-world data stores to improve AI performance.

New AI Tutor Achieves 0.71-1.30 SD Effect Size In Dartmouth Course [Pdf]

A new AI tutoring system achieved effect sizes of 0.71 to 1.30 SD in Dartmouth’s course, marking a notable advance in educational technology.

Astrophysicists Puzzle Over Webb’s New Universe

Scientists are analyzing surprising observations from the James Webb Space Telescope that challenge current understanding of the universe.

More Than Half Of Adults In U.S. Say They Lack Basic Statistical Understanding

A new survey reveals over 50% of American adults struggle with fundamental statistical concepts, raising concerns about public understanding of data.