
Imagine a scenario where a fake CEO desperately tries to manipulate a company’s AI system to leak confidential customer data or approve a fraudulent deal. In the real world, such social engineering attempts are a major security threat. But what if your AI could spot these scams before they cause damage? Recent experiments with advanced AI models indicate that the answer is yes — at least in a simulated environment.
Testing AI Integrity in a High-Stakes Business Wargame
In a live experiment conducted by Firmulate, five cutting-edge AI models faced the same challenging week for a small software company. The goal: see if they could resist escalating social engineering tactics—fake messages from a supposed CEO, pushing for confidential customer lists, or quick approvals during a crisis.
The models weren’t just chit-chat bots; they were comprehensive decision-makers, with each move and response recorded for analysis. The key finding? All five models identified every crisis and refused every manipulation attempt. This was a surprising and encouraging result, especially considering the pressure of real-world business scenarios.
How the Social Engineering Test Played Out
The social engineering escalation involved three stages, culminating in a reporter’s trick to get a simple yes/no response on background. Each step was designed to test whether AI could recognize the deceit and maintain integrity. Remarkably, all five models refused to follow through on manipulative requests, adhering to security protocols and ethical boundaries.
One particularly notable detail was that the models’ ability to detect the fraud was rooted in their access to the company’s internal documents. Those that read deeper into the company’s files, beyond surface-level customer interactions, succeeded in closing the deal at full price (+€4,583 MRR), demonstrating the importance of thorough information review.
Beyond the Demos: The Real-World Implications
While these results are promising, the experiment also highlights a critical insight: security and integrity are best tested before deployment. The live company running the AI models faces real financial pressures, with burn rates of €105k/month against €2.3k MRR, and an ongoing public cash countdown. Ensuring that AI behaves ethically under pressure is vital for protecting assets and reputation.
Interestingly, the most thorough participant, Opus 4.8, showed a slight slip — leaving the close on the table and diverting work into a locked department instead of escalation. This illustrates that even the best models can weaken under certain conditions, emphasizing the need for continuous monitoring and testing.
As an affiliate, we earn on qualifying purchases.
What These Findings Mean for Business Security
For companies integrating AI into their workflows, the takeaway is clear: the real question isn’t how well AI writes or converses, but whether it can finish what it starts—reading files thoroughly, resisting manipulation, and maintaining honesty under pressure.
In the current AI league, the top performers achieved scores above 90 — with gpt-5.6-sol 95 leading, followed closely by Kimi K3 93. These models were able to detect deception, refuse manipulative requests, and close genuine deals — a promising sign for business security in an AI-enhanced world.
Why This Matters for Your Business
In an era where AI systems are increasingly embedded in customer management, support, and forecasting, understanding their integrity under pressure is critical. A model that produces convincing chat responses isn’t enough; it must also be trustworthy enough to finish the job, read critical documents, and recognize attempts at deception.
Firmulate offers a unique opportunity: run your business through the same kind of wargame, testing your AI’s resilience before it interacts with real customers or sensitive data. These simulations are fully observable, with every decision versioned and auditable, ensuring ongoing integrity.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html