
Imagine running a real company that loses €105,000 each month, yet you can watch its every move unfold live online. This is not fiction — it’s the extreme experiment of a software startup using artificial intelligence as its core decision-maker, publicly battling to stay afloat while exposing its weaknesses and strengths in real time.
The Experiment: AI as a Company Manager in the Wild
At the heart of this daring project is Firmulate, a real, functioning software company with 13 synthetic employees driven by AI models. Every business day, the company faces simulated crises, market temptations, and decision-making puzzles. What makes this experiment extraordinary is that it’s public — anyone can follow along at firmulate.com/live.html.
Each AI model is tested under identical conditions: same clients, same challenges, and the same pressure to cheat or manipulate. The models are judged solely on their ability to navigate these scenarios ethically and effectively, with decisions recorded and auditable for analysis.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Results That Speak Volumes
The findings are striking:
- All four AI models identified every crisis and refused every manipulation attempt, demonstrating a strong capacity for integrity and crisis awareness.
- Only two of these models managed to sign a lucrative €55,000 deal after analyzing the company’s own internal files — a critical detail buried within documents that most human decision-makers might overlook. This buried fact was key to closing the deal, adding €4,583 monthly recurring revenue (MRR).
- The models that read the files won the deal at full price, revealing the importance of internal knowledge in decision-making.

Responsible AI in the Age of Generative Models: Governance, Ethics and Risk Management (Byte-sized Learning)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Confronting Social Engineering and Ethical Challenges
The experiment also tested how AI handles social engineering tactics, such as fake CEO messages escalating through stages and even a reporter trick. Remarkably, all five models refused to be manipulated, with Kimi K3 explicitly treating such requests as potential impersonation or approval bypass attempts.

AI for Public Relations: A How-To Guide for Implementation and Management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Reality of a Burnt-Red Business
This live experiment isn’t just a theoretical showcase — it’s a real business that’s losing money every day. The company burns through €105,000 each month against a modest €2,300 in monthly recurring revenue. Despite the financial strain, it operates transparently; the entire process is viewable online, with decisions, strategies, and even internal conflicts on display at firmulate.com/live.html.
The company’s management decisions are also put to a test through a public quiz, where 242 real decisions are analyzed, and viewers can try to guess which AI model made each call at firmulate.com/quiz.html.
As an affiliate, we earn on qualifying purchases.
Deep Analysis, Flaws, and Lessons
Among the four models tested, Opus 4.8, with the deepest rule set (over 80 learned rules), performed the most thoroughly. Yet, it still left a deal on the table due to disciplined lapses like failing to escalate issues and instead writing attempts into a locked department.
Interestingly, the Kimi K3 model, which ran with default settings, showed the cleanest discipline and was the only one to sign the deal at full price. This suggests that less aggressive or default configurations might foster better integrity, even if they score slightly lower on other metrics.
Implications for Business and AI
This experiment raises profound questions: as AI agents become more involved in critical business functions — from customer support to strategic decision-making — the real concern isn’t just whether they produce coherent text. It’s whether they stay honest, follow internal protocols, and complete tasks reliably under pressure.
The experiment underscores that in high-stakes environments, the ability to read internal files, resist manipulation, and execute decisions ethically is just as crucial as their technical capabilities.
The Build-in-Public Innovation
What sets this apart is its transparency. The company publishes every decision, every version, and every rule learned, effectively building a real-time case study accessible to anyone interested in AI, ethics, and business resilience. It’s a form of build-in-public that exposes the raw truths of managing AI-driven companies in volatile environments.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html