AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

As AI models become more integrated into business operations, understanding how they perform under real-world pressures is crucial. But what does it really mean when an AI earns a baseline score of 26 out of 100, even when doing nothing?

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

Unpacking the Benchmark: More Than Just a Score

At first glance, one might assume a “do-nothing” AI baseline should score zero. Surprisingly, it scores 26 points in the latest crucible league—a comprehensive test that simulates the worst week for a small software company. Why? Because this baseline accounts for partial progress, recognizing even minimal efforts, but also incorporates strict rules that cap the total score if trust is breached.

The Methodology Behind the Score

The experiment involved running several frontier AI models through identical scenarios: same customers, same crises, same manipulations. Every decision was recorded, versioned, and auditable. The models’ ability to navigate crises, resist manipulation, and maintain integrity was meticulously assessed.

Why Partial Progress Matters

Interestingly, all models—regardless of their overall scores—detected every crisis and refused every manipulation attempt. Yet only two managed to finalize and sign a €55,000 deal based on their own analysis. The others, despite identifying opportunities and resisting deceit, failed to follow through to closure, highlighting the importance of execution and discipline in AI decision-making.

Amazon

AI model testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness: Reading Deeper into Documents

One of the key findings was that the decisive advantage often lay not in surface-level responses but in reading and understanding documents two references deep within the company’s files. Those models that successfully read and interpret these internal documents achieved the full deal value, worth over €4,580 in monthly recurring revenue.

Trust and Integrity Under Pressure

Models faced social engineering tactics, including fake CEO messages and reporter tricks designed to escalate requests over multiple stages. All five models refused to participate in these manipulations, citing reasons like suspicion of impersonation or approval bypass. This demonstrates that AI can be trained to prioritize trustworthiness over mere task completion.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Live Experiment: Real Company, Real Money, Real Risks

Firmulate runs an ongoing live test with a simulated company employing 13 synthetic employees handling real business mechanics—burning €105,000 each month against a revenue of €2,300. The environment is fully transparent and observable, with every decision versioned, allowing for continuous assessment of management quality.

The Performance Spectrum

Among the tested models, Opus 4.8, the most thorough participant with over 80 learned rules and deep analyses, unfortunately finished last in the critical deal, leaving the close on the table due to slips in discipline. This highlights that thoroughness alone doesn’t guarantee success—execution matters just as much.

Amazon

AI trustworthiness evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Implications for Business and Education

This experiment underscores key lessons for both sectors. For business, it emphasizes the importance of transparency, trust, and discipline in AI decision-making—factors often overlooked in superficial evaluations. For educators, it offers a tangible example of how AI models are tested under real-world pressures, illustrating the importance of ethical behavior, thorough understanding, and execution in AI deployment.

The Takeaway

What sets this benchmark apart is its honesty. It recognizes that partial progress counts but also that a breach of trust caps the score. In other words, progress is valuable, but integrity is non-negotiable. For organizations considering AI integration, these findings highlight that performance under pressure, trustworthiness, and the ability to read and interpret deeply embedded information are critical success factors—and metrics that matter far more than flashy chat demos.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

AI benchmarks that reflect real-world pressures, trustworthiness, and discipline provide more meaningful insights for businesses and educators alike. Progress counts, but integrity is paramount in AI performance.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

business AI simulation platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Models Show Resilience Against Social Engineering Attacks in Live Company Test

A recent live experiment shows that five AI models refused social engineering tricks during simulated crises, proving that integrity can be embedded before deployment.

Some Combinatorial Applications Of Spacefilling Curves

Researchers have demonstrated new combinatorial uses of space-filling curves, opening potential for optimized data structures and spatial algorithms.

How to Choose Scientific Calculator For Students

Step-by-step guide for students to effectively operate a scientific calculator for math and science tasks, ensuring accurate results.

Inside a Reckless AI-Driven Business That’s Publicly Fighting to Survive

A real, losing company publicly tests AI models under crises, revealing that honesty, internal knowledge, and resilience are key to effective AI-driven decision-making.