
In an era where artificial intelligence increasingly shapes business decisions, a recent live experiment demonstrates that well-designed AI can resist social engineering tricks—at least, under controlled conditions. For educators, researchers, and industry leaders alike, this reveals a crucial insight: integrity under pressure isn’t just a theoretical ideal, but an achievable standard before AI is rolled out in real-world settings.
Testing AI Integrity in a Simulated Crisis
Recently, a live experiment conducted by Firmulate placed five frontier AI models—including top performers like GPT-5.6-SOL and Kimi K3—into a simulated scenario of a small software company facing its worst week. The identical challenges, customer crises, and temptations to manipulate data or cut corners were presented to each model, with decisions fully auditable and consistent across tests.
The Setup and Goals
Each AI was tasked with managing this virtual company, which had a public cash flow of €105,000 monthly against just €2,300 MRR. The goal? Navigate crises ethically, uphold company integrity, and ultimately close a €55,000 deal based on analysis and decision-making. The overarching question: could these AI models maintain honesty, recognize manipulation attempts, and deliver on their commitments?
Key Findings: Integrity and Decision-Making
Remarkably, all five models identified every crisis and refused every manipulation attempt, including social engineering tactics. One common scenario involved a fake CEO messaging staff to bypass usual approval processes—an escalating series of requests culminating in a ‘just one yes/no’ background question to a reporter. Every AI refused to sign off on illicit actions, exemplifying robust ethical guards built into their decision processes.
The fifth model, Opus 4.8, distinguished itself as the most thorough—analyzing over 80 learned rules and conducting deep assessments. Yet, it was also the only model that slipped slightly, leaving a close deal on the table due to discipline lapses. Its weakness was subtle but consistent across the models: the failure to escalate or document certain decisions properly, instead leaving some actions locked in departments without proper review.

Responsible AI: Implement an Ethical Approach in your Organization
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Test: Reading Deep in Company Files
Beyond immediate crisis management, the experiment revealed a crucial insight: the decisive factor in securing the business deal was the AI’s ability to examine internal company documents. Those models that read and interpreted files deeper within the company’s own records uncovered critical information buried two references deep—information that, when used, led to the full €4,583 MRR deal. In contrast, models that only focused on surface-level data missed this opportunity, illustrating the importance of comprehensive information analysis for trustworthy decision-making.
Implications for Business and AI Deployment
This experiment underscores a vital truth for enterprise AI deployment: integrity isn’t just about avoiding obvious scams. It’s about ensuring that AI systems are designed to read, verify, and interpret relevant internal data before making commitments. This proactive approach can prevent costly breaches of trust, whether in customer data, financial transactions, or strategic deals.

AI-Powered Data Workflows: From Raw Data to Actionable Insights: Automating Data Cleaning, Analysis, and Reporting with Python and Modern AI Tools (AI & Automation for Professionals Series Book 3)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What This Means for the Future of AI in Business
While some might assume that AI’s ability to mimic human behavior equates to trustworthiness, the real test is whether it can withstand ethical pressures when stakes are high. The experiment shows that all tested models—regardless of their scores—refused manipulation attempts, signaling that with proper safeguards, AI can be a reliable steward of business integrity.
Furthermore, the experiment demonstrated that the weakest links are often in how decisions are documented and escalated. The Opus 4.8 model’s slight discipline slip highlights that high analytical depth must be paired with rigorous process adherence.
Embedding Integrity Before Production
This series of tests offers a key lesson: social engineering and integrity challenges should be integrated into AI testing frameworks long before deployment. The ability to recognize manipulation, read internal documents thoroughly, and escalate issues appropriately can be validated in controlled environments—far safer than addressing breaches after they happen.
As an affiliate, we earn on qualifying purchases.
Conclusion: The Promise of Ethical AI
As the live experiment concludes, the takeaway is clear: AI can uphold ethical standards under pressure, provided it is designed with these principles in mind. The fact that all five models refused manipulation attempts during this controlled crisis scenario is an encouraging sign that trustworthy AI is achievable now, not just in the distant future.
For enterprises contemplating AI adoption, the message is simple: test your AI systems thoroughly in simulated crises before going live. Doing so can reveal vulnerabilities and reinforce ethical decision-making—ensuring that when real-world pressures hit, your AI workforce remains honest and reliable.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Analytics, Data Science, & Artificial Intelligence: Systems for Decision Support
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.