AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get pool and patio gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

When the pool business hits rough water, what would your AI do?

A burst pipe, a sudden customer exodus, a competitor undercutting your prices: for a pool service or patio company, a difficult week can put customer trust and cash flow under pressure at the same time. Firmulate’s live experiment asks a bigger question than whether an AI can write a convincing email: can it run a company, make disciplined decisions and close a deal when the stakes are real?

The experiment is public and watchable at Firmulate. Its final Crucible League results offer a glimpse of what AI management can get right—and where it may still leave opportunity on the table.

One company, the same hard week

In the July 2026 final, each frontier model faced the same small software company, the same customers, the same crises and the same temptations. Decisions were versioned and auditable. The live company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. Its workdays are versioned, and its playbook has learned more than 680 rules.

The point is not to mistake a simulation for a pool-service business. It is to watch models handle pressure before trusting AI with decisions that touch customers, pricing or operations. The league ranked gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

Seeing the problem was not the same as finishing the job

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s blunt summary: “Same diagnosis, same pitch — no signature.” A model can sound sure about the right move and still fail to make it.

The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. For a pool company, the lesson is practical: useful AI needs to find the relevant details in company records and carry a sound recommendation through to action.

Trust and discipline under pressure

In staged social-engineering attempts, fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s reasoning was to “Treat the request as a suspected approval-bypass / possible impersonation.” That kind of skepticism matters wherever an AI assistant might encounter requests involving customer information, money or company authority.

The bottom-placed participant, Opus 4.8, was the most thorough: it learned 80 rules and produced the deepest analyses. But the deal was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models. K3 also ran without an effort parameter, using the API default, while the others ran at xhigh—a fairness detail to keep in mind when reading the ranking.

Firmulate’s live site also offers a “guess the model” quiz built from 242 real, unedited management decisions. Readers can see how the choices played out before deciding which model they think made them.

From watching to trying it against your business

The experiment’s next step is a pilot for enterprises. It uses a read-only export of a company’s own business to create a digital twin, then runs crisis scenarios against it and produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. For a pool, patio or water-lifestyle business, that offers a way to examine how AI might handle your own customer and operational pressures before letting it near live workflows.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Try the wargame on your own terms

The live league shows that spotting a crisis, protecting trust and completing the job are separate tests. A pilot lets enterprises examine those tests against their own business using a read-only export. To discuss a pilot, visit Firmulate’s pilot page or email contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Heater Type That Fits Windy Patios Better Than Calm Ones

Irdnfrared heaters are ideal for windy patios, offering consistent warmth despite gusts, and here’s why they outperform traditional options in outdoor conditions.

Summer Grilling Hack: Mastering the Ninja Foodi Smart XL Air Fry Oven

Discover tips and hacks for using the Ninja Foodi Smart XL Air Fry Oven to elevate your summer meals with crispiness and convenience.

Why Outdoor Privacy Screens Improve More Than Just Appearance

Creative outdoor privacy screens enhance your space by offering more than looks—they provide comfort, protection, and serenity that transform your backyard experience.

The Best Way to Plan a Backyard Movie Night Without Technical Stress

Just follow these simple steps to create a stress-free backyard movie night and enjoy a memorable evening under the stars.