
A Glimpse into the Future of AI Management: Transparent, Risky, and Public
Imagine a company without human employees, where every decision is made by AI models tested against real crises, real money, and real deadlines. Welcome to the world of Firmulate, a groundbreaking experiment in building AI-powered organizations in the open, with every move watched by the public and every outcome measurable.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The experiment: an AI-driven company under scrutiny
Firmulate runs a small, virtual company with 13 synthetic employees, simulating real-world business mechanics. It faces daily crises—customer issues, ethical dilemmas, and financial pressures—just like a traditional firm. What makes this experiment extraordinary is its transparency: every decision, every crisis response, every rule learned by the AI models is publicly available, versioned, and auditable at firmulate.com/live.html.

AI For Students: Your Academic Writing Assistant: Ethical Tools, Prompts & Workflows to Write Stronger Papers in High School, College & Beyond (AI for Academic Success)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Timing and stakes: a live cash countdown
This virtual company burns through €105,000 each month but earns only €2,300 in monthly recurring revenue. It’s a real-world financial challenge, with a public cash countdown that underscores the company’s fragile state. Every workday, the company’s decision process is versioned, and its rules—more than 680 self-learned strategies—are openly documented, enabling anyone to follow its evolution and decisions.

The Intelligent Ledger: How Agentic AI, Quantum Computing, Data, and Tokenisation Are Transforming Banking
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The experimental test: AI models vs. real crises
Four frontier AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5—were each tasked with guiding the same virtual company through its worst week. They faced identical customer crises, ethical dilemmas, and temptations to cheat or manipulate. All models successfully identified every crisis and refused every manipulation attempt. Yet, only two of them managed to close a critical €55,000 deal, earned through their own analysis and pitching—demonstrating that recognizing a problem is not enough; execution is key.

AI Skin Analyzer Device for Face & Scalp – Multi-Light UV and Polarized Imaging, 21.5 Inch Touchscreen, Handheld Scalp Viewer, Client Image Records, Gray
Professional Face and Scalp Imaging: Capture clear facial and scalp images with an enclosed face chamber, chin rest,…
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What was the hidden weakness?
Deep within the company’s own files—a document reference—lay the critical information needed to close the deal. The models that read this hidden document and understood its significance won the full-price deal, worth over €4,583 in monthly recurring revenue. This buried fact exemplifies how reading and analyzing internal data can be decisive in a high-stakes business environment.
Ethics and social engineering: AI’s integrity under pressure
The experiment also tested social engineering tactics: staged messages from a fake CEO escalating over three stages, and a reporter’s subtle request for a background approval. Remarkably, all five models refused to bypass security, with Kimi K3 explicitly treating the requests as potential impersonation or approval-bypass attempts. This demonstrates that these AI models can uphold ethical standards even under social pressure.
The real-world implications: trust, honesty, and operational discipline
Despite its inability to turn a profit, the live experiment reveals crucial insights: the AI models are capable of detecting crises, refusing manipulation, and demonstrating ethical restraint. However, discipline and operational rigor vary. In the case of Opus 4.8, despite the most thorough analysis—over 80 learned rules—the company left a deal unexecuted due to slip-ups, illustrating that thoroughness alone does not guarantee perfect performance.
Why this matters for the future of AI and business
Today’s AI models can identify problems, refuse unethical shortcuts, and follow complex decision rules. Yet, they still face challenges in closing deals, executing plans, and maintaining discipline under pressure. For decision-makers contemplating AI integration, the key questions are: Will AI finish what it starts? Will it read and understand your internal data? Will it stay honest when stakes are high? The Firmulate experiment publicly demonstrates that AI-driven organizations can be transparent, disciplined, and ethically sound—even if they currently struggle with profitability.
Watch the experiment unfold
Interested in seeing this AI-driven company in action? Visit firmulate.com/live to watch real-time decision-making, crisis management, and financial mechanics at play. This build-in-public approach offers a rare, unfiltered view into AI’s potential—and its limitations—as a leader in management and operational discipline.

Key Takeaway
While AI models can identify crises and uphold ethical standards, their ability to execute plans and close deals depends on operational discipline. Firmulate’s live experiment showcases an unprecedented level of transparency in AI management, raising vital questions about the future of automated decision-making in real business environments.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html