
Imagine running a business in full view of the world—making daily decisions, facing crises, and watching your financial health plummet in real time. For a pioneering AI experiment, this is not science fiction but a reality, offering a rare glimpse into how AI performs as an actual company—warts and all.
The Live Experiment: An AI-Run Company in Action
At the heart of this experiment is a live digital company run by 13 synthetic employees, with real money mechanics in play. Every workday, the company faces genuine crises—ranging from customer issues to operational temptations—and each decision is meticulously recorded and versioned for public scrutiny. This setup isn’t just a test; it’s a transparent, ongoing experiment in AI management and decision-making.
Currently, the company is losing €105,000 each month against a modest €2,300 monthly recurring revenue (MRR). Despite the financial loss, the experiment offers invaluable insights into how different AI models handle complex, real-world business scenarios. You can watch the company in action at firmulate.com/live.html.

Decision Making Under Uncertainty: Theory and Application (MIT Lincoln Laboratory Series)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The AI Models: Competing at the Frontiers of Decision-Making
Four frontier AI models were tasked with navigating this challenging environment, each undergoing the same difficult week of decisions, crises, and ethical tests. These models include:
- gpt-5.6-sol 95 (top scoring)
- Kimi K3 93 (newcomer with highest discipline)
- Sonnet 5 88
- Fable 5 77
All four successfully identified every crisis and refused manipulation attempts—such as fake CEO messages or reporter tricks. Notably, only two of these models managed to close the €55,000 deal their own analysis indicated they could secure. The difference was in how they read and interpreted critical company documents; models that read deeper into the company files discovered a hidden, decisive piece of information that led to the deal, earning an additional €4,583 MRR.
Lessons on Trust, Ethics, and Performance
This experiment underscores a fundamental truth: AI’s ability to navigate complex business ethics and be trustworthy under pressure isn’t about writing pretty chat responses. It’s about consistency, thoroughness, and integrity in decision-making. The models were tested against social engineering tactics, like staged CEO messages, and uniformly refused to be manipulated—showing a strong grasp of ethical boundaries.
Interestingly, the most thorough participant, Opus 4.8, often provided deeper analysis, learned over 80 rules, but still left some opportunities on the table, like not escalating issues appropriately. This highlights that even the most advanced models are still imperfect, especially when discipline wanes under stress.
The Broader Significance for Business and AI
This ongoing company is not just a curiosity; it’s a mirror reflecting how AI might perform in real business environments. As firms consider deploying AI agents in customer relations, support, or forecasting, the questions aren’t merely about language fluency but about the AI’s ability to finish what it starts, read relevant information thoroughly, and maintain ethical standards under pressure.
For those interested, the experiment is publicly available and fully transparent. Every decision, every crisis, every ethical challenge is recorded and open for scrutiny, providing a rare look into AI’s actual capabilities and limitations—beyond the hype of chatbots and demos.
Why This Matters to You
If you manage a pool, patio, or water feature business, you might wonder how AI could change your industry. The insights from this experiment show that AI is already testing itself in complex, high-stakes environments. The key takeaway: the real value of AI isn’t just in generating responses or ideas, but in executing tasks reliably, ethically, and consistently—especially when the pressure is on.
As the live company continues to burn through cash and make tough decisions in the open, it exemplifies what building in public truly means: transparency, accountability, and the relentless pursuit of improvement. Watching this unfold offers a rare educational opportunity for any business contemplating AI adoption today.
Next Steps: Engaging with the Experiment
Curious to see how AI manages real-world crises or to test your own business decisions against these models? You can explore the live experiment, read detailed quotes from the models, participate in quizzes, or even run your own business wargame at firmulate.com/quiz.html and firmulate.com/pilot.html. This is building-in-public at its most transparent and instructive.
In a world where AI’s role is rapidly expanding, observing how it handles the messy realities of business—losing money, facing crises, and resisting manipulation—could be the most valuable lesson you take away today.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html