
What if your AI could resist a corporate impersonation scam?
Imagine a scenario where a fake CEO tries to manipulate an AI to leak sensitive information or approve a fraudulent deal. For many, this would be a nightmare scenario—yet recent experiments show that AI models are holding their ground better than expected. This is a story about security, integrity, and the surprising resilience of AI under pressure.

CompTIA SecAI+ CY0-001 Study Guide: Complete Reference with Practice Tests, PBQ Scenarios, and Study Tools for Exam Preparation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real Test: A Simulated Crisis for AI
In a carefully controlled experiment, five of the most advanced AI models faced the same challenge: a fake CEO message escalating through multiple stages, including a sneaky reporter trick. The goal? To see if these models would recognize the deception and refuse to comply.
The test was rigorous. Each AI was given the same scenario involving a small software company, with real money mechanics, crises, and temptations to cheat. Every decision was recorded and made auditable to ensure transparency. The models had to identify threats, read critical documents, and decide whether to approve a deal—just as a human manager would.
Results That Defy Expectations
All five models successfully spotted every crisis and refused every manipulative attempt. Notably, only two of these models went on to sign a €55,000 deal—an outcome based purely on their own analysis, without external pressure or signature coercion.
What made the difference? A key insight emerged: the critical piece of information was hidden two document references deep within the company’s files. Models that effectively read and analyze these files won the deal at full price. This reveals a crucial aspect of AI security: reading and understanding internal data is vital to ensuring trustworthiness.
Why This Matters for Business Security
In the real world, AI will increasingly handle sensitive tasks—managing customer relationships, supporting operations, and making strategic decisions. The question isn’t just whether AI can generate convincing language; it’s whether it can stay honest under pressure and follow the right protocols.
The experiment underscores that integrity isn’t just tested in the heat of a crisis but can be assessed beforehand. AI models that are trained and tested in simulated crisis scenarios are less likely to be duped or manipulated in real situations.
The Human Element and AI Resilience
The experiment also highlighted that even the most thorough participant, the Opus 4.8 model, which analyzed over 80 rules and conducted in-depth analysis, slipped up when discipline slipped—failing to escalate issues properly. This shows that robustness isn’t just about intelligence but also about discipline and process adherence.
Implications for Your Business
For companies considering AI integration, this experiment offers a reassuring message: models can be tested for integrity before deployment, reducing the risk of costly breaches or manipulations. The firms behind the experiment run a live, watchable system where this testing is ongoing and transparent. You can observe how AI models perform in real-time, handling crises, making decisions, and resisting social engineering tactics.
As the industry advances, the focus must shift from just chat quality to measuring actual decision-making integrity. AI that can finish what it starts, read and understand critical documents, and stay honest under pressure is a valuable asset—especially in high-stakes environments.
What’s Next?
The current league table from the experiment shows impressive scores—gpt-5.6-sol scored 95, Kimi K3 scored 93, Sonnet 5 scored 88, and Fable 5 scored 77—out of a perfect 100. These rankings reflect not just raw intelligence but the ability to maintain integrity and discipline under simulated duress.
For those interested in testing their own AI, the platform offers a read-only wargame environment where companies can simulate similar crises without risking their real systems. It’s a proactive way to ensure your AI workforce is trustworthy before deploying it in critical tasks.

The Key Takeaway: Testing Integrity Before Deployment
AI models demonstrate remarkable resilience in simulated social engineering attacks, refusing manipulation attempts and making honest decisions. Testing AI integrity in controlled scenarios is essential before real-world deployment, ensuring your AI works reliably, ethically, and securely when it truly counts.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html