
Imagine a world where artificial intelligence doesn’t just assist with tasks but actually runs a company—making decisions, managing crises, and even sealing deals. While AI’s capabilities often focus on language and pattern recognition, recent experiments reveal its management personality traits, including honesty, discipline, and risk appetite. For educators, scientists, and decision-makers alike, understanding these AI ‘personalities’ can be crucial when integrating automation into real-world operations.
Testing AI Decision-Making in a Real-World Business Simulation
Recently, a groundbreaking live experiment placed four advanced frontier AI models in the role of a small software company’s management team. The goal? To navigate the company’s worst week—handling customer crises, internal conflicts, and ethical temptations—while decisions were recorded, verified, and analyzed.
All four models faced identical scenarios, including a fake CEO message escalating tensions and a reporter asking for a discreet yes/no approval. They were monitored in a fully transparent environment, where every choice could be reviewed and scored, revealing not just their decision but the reasoning behind it.
AI decision-making management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Findings: Integrity, Disciplined Decision-Making, and the Power of Information
Remarkably, every AI model identified every crisis and refused every attempt to manipulate their decisions. They maintained integrity under pressure, a crucial trait for automation systems that could impact financial or operational outcomes. However, when it came to closing deals, their behaviors diverged—highlighting distinct management personalities.
The most successful was GPT-5.6-sol, which identified a buried factual reference in internal files—something that other models overlooked—and used that information to close a €55,000 deal. This model scored a perfect 95 out of 100, demonstrating thorough analysis and honest commitment.
Close behind was Kimi K3, scoring 93. It also secured the deal, but with a slightly cleaner discipline record—refusing manipulative requests and sticking to the facts. Conversely, Sonnet 5 scored 88, closing the deal but slipping on some procedural discipline, while another iteration of Sonnet scored 77, missing opportunities and leaving the deal unsealed.
As an affiliate, we earn on qualifying purchases.
Understanding the AI Personalities: Who Stays Honest and Who Slips?
One revealing aspect was that the decisive weaknesses often resided in internal document analysis, not in customer interactions. The models that read deeper into company files, like GPT-5.6-sol, won at full price, while others failed to leverage internal knowledge effectively. This suggests that a model’s ability to access and interpret relevant internal data correlates strongly with successful, honest decision-making.
Moreover, in social engineering tests—fake corporate messages and reporter tricks—all models refused to cooperate, citing suspicion and impersonation concerns. Kimi K3 explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” This indicates a built-in risk-averse stance and cautious management style.
AI internal data analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for Business and AI Adoption
What does this mean for organizations contemplating AI automation? First, that models can be trusted to recognize crises and resist manipulation. Second, that their management style—whether rigorous, cautious, thorough, or lenient—varies significantly, impacting outcomes like deal closures or procedural discipline.
The experiment was conducted in a live, operational environment with 13 synthetic employees and real money mechanics—burning €105,000 monthly against a modest €2,300 recurring revenue. The system is fully transparent and auditable, showing that these models are more than just chatbots—they are active participants in complex management tasks.
AI ethical decision support system
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps: Wargaming Your AI Workforce Before Deployment
Businesses can try this approach with their own systems through Firmulate’s interactive quiz. The platform allows companies to simulate management crises with their AI agents, ensuring the models exhibit the desired traits before real-world deployment. It’s a way to test whether your AI will finish what it starts, stay honest under pressure, and read the critical internal data—traits vital for responsible automation.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html