
Imagine managing a busy poolside business where every decision counts — from handling customer complaints to sealing big deals. Now imagine trusting an AI to run that business for a week. Would it stay honest? Would it get the job done? These are the questions behind a groundbreaking live experiment that pits the latest AI models against real-world management challenges.
The Experiment: Putting AI in the Director’s Chair
Today, AI isn’t just about chatbots or recommendation engines — it’s increasingly being tasked with complex management decisions. To test just how reliable these models really are, a company has set up a live, ongoing simulation where four top frontier AI models run a small software business through its most tumultuous week.
This isn’t a game or a demo; it’s a real-time, auditable battle of decision-making under pressure, with real money at stake. The company, which operates with 13 synthetic employees and a range of real money mechanics, faces identical crises, customer demands, and temptations — only the AI model changes from run to run.
AI management decision support tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How the Models Were Tested
Each AI was given the same set of challenges and had to make decisions that could impact the company’s bottom line and integrity. Their performance was scored based on their ability to identify critical information, refuse manipulative requests, and close profitable deals. Crucially, the models were also tested for honesty — whether they would fall for social engineering, such as fake CEO messages or reporter tricks.
All four models demonstrated impressive crisis awareness and refused every manipulation attempt. Yet, only half of them managed to close the most valuable deal — a €55,000 contract worth +€4,583 MRR. The other two models, despite similar diagnoses, failed to secure the deal, leaving potential revenue on the table.
What Made the Difference?
The key to success was what was buried deep within the company’s files. The models that successfully closed the deal read two document references that contained crucial information about a competitor weakness. Those that missed this detail failed to act on the full context, losing out on a substantial opportunity.
Interestingly, all models detected every crisis and refused manipulative social engineering — even escalating fake CEO messages in a three-stage process plus a reporter trick. In the words of Kimi K3, which performed well in fairness, “Treat the request as a suspected approval-bypass / possible impersonation.” This shows that these models are capable of recognizing and refusing deceptive tactics, a vital trait for trustworthy AI management.
The Real-World Implications
This experiment isn’t just about software games; it’s a window into how AI could manage real companies, or support human managers, in high-stakes situations. The live company runs daily, with over 680 self-learned rules, and faces ongoing financial challenges — burning €105k a month against a mere €2.3k MRR. Watching this in real time reveals how different AI models behave, adapt, and sometimes slip in discipline or focus.
Among the tested models, Opus 4.8 was the most thorough, analyzing over 80 learned rules and providing deep insights. However, it also left opportunities unseized, like the close deal that the other models secured. Conversely, the least disciplined model, despite thorough analysis, struggled to close deals and stayed inside its boundaries, missing bigger gains.
Measuring AI Personalities — It’s Not Just About Scores
The performance scores from the final leaderboard were telling. GPT-5.6 scored 95 out of 100, having identified the hidden critical fact and closing the deal. Kimi K3 scored 93, demonstrating the cleanest discipline and closing the same deal. Sonnet 5 scored 88, and another Sonnet model scored 77, both closing deals but with more slips.
These scores reflect more than raw decision-making; they hint at different management personalities embedded within each model. Some are thorough and cautious, others disciplined but less aggressive. Notably, K3 ran without an effort parameter, making its behavior more fair and controlled, while others ran at a high effort setting, pushing for maximum output.
Why Should Pools & Water Lifestyle Enthusiasts Care?
If you manage a pool or patio business, you’re already balancing multiple priorities — customer satisfaction, safety, and revenue. You might wonder: Could an AI handle these chores reliably? The answer, demonstrated here in a real live environment, is that AI can be trustworthy enough to make critical decisions, read vital data, and defend against manipulation. But it’s also clear that not all AI models are created equal.
Before trusting AI with your business, you need to test and understand what kind of ‘manager’ it would be. Firmulate’s live experiments let enterprises run a simulated, risk-free version of their operations, revealing how each AI would perform under pressure. You can see firsthand which model is disciplined enough to stay honest, focused enough to seize opportunities, and smart enough to avoid costly slips.
The Final Word
With AI becoming a bigger part of business operations—from CRM to support and forecasting—it’s essential to ensure it can finish what it starts. This live experiment is a valuable step in understanding AI’s management personality, its trustworthiness, and its potential to replace or support human decision-making.
Curious to see how your AI workforce might perform? Try the same test with your own business data at firmulate.com/quiz.html and discover which model aligns best with your management style.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html