AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get pool and patio gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

What if your next AI assistant could run your business — and actually finish what it starts?

Imagine deploying an AI that’s not just good at chatting but can handle your company’s toughest crises, read your documents thoroughly, and resist shady manipulations. That’s the promise—and the challenge—facing AI developers today. Recent live experiments reveal that some models are starting to deliver on this potential, outperforming others in real-world tests that mirror the chaos of actual business operations.

Amazon

AI business management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Live Experiment: Putting AI to the Test in a Real Business Setting

In a groundbreaking live experiment conducted by Firmulate, four frontier AI models were tasked with running a small software company through its most challenging week. This wasn’t a simulation in a chat interface; it was a real-time, auditable test that involved actual crises, customer decisions, and the temptation to cheat. The goal: see which AI could truly manage the company’s worst week — not just talk about it.

All four models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5—demonstrated a strong ability to recognize crises and refused every manipulation attempt, including social engineering tactics like fake CEO messages. Yet, the differences lay in their ability to complete critical tasks and close deals. Kimi K3, a newcomer from Moonshot, scored a 93 out of 100, just behind the top scorer, gpt-5.6-sol, with 95. The standout for K3 was its ability to find a buried security fact deep within the company’s files—a crucial detail that clinched a €55,000 deal, translating into +€4,583 monthly recurring revenue (MRR).

Cracking the Hidden Code: Kimi K3’s Edge

The experiment uncovered a buried weakness common among models: the failure to read and interpret key documents. While the top models identified critical facts buried two document references deep, others left the deal on the table or slipped discipline, showing weaker performance. K3’s success highlights an important insight: reading comprehension at depth can be decisive in real business outcomes, not just in chat quality.

Honesty and Integrity Under Pressure

The experiment also tested models against social engineering—fake CEO messages escalating in three stages, plus a reporter’s trick question. All models refused to approve or sign any fake requests, citing suspicion and adherence to protocol. Kimi K3’s on-record reasoning was clear: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates an encouraging level of honesty under pressure, a crucial trait for AI in business applications.

The Real Business Environment

The live setting involved 13 synthetic employees, real cash mechanics, and a public focus on daily performance metrics. The company burned €105,000 monthly against only €2,300 in MRR, with a public cash countdown actively ticking down. Every decision was versioned and auditable, and the entire process is observable at firmulate.com/live. This transparency underscores how AI models perform outside of scripted demos, facing genuine operational dilemmas.

The Lessons for Business Leaders

While the most thorough participant, Opus 4.8, scored the lowest (73), it shed light on discipline weaknesses, such as leaving deals on the table or failing to escalate issues. The results emphasize a vital point: the best AI models are those that not only recognize crises but also follow through consistently and honestly, especially under pressure.

Fairness Note and the Open League

It’s important to note that Kimi K3 ran without an effort parameter (the API default), while the other models ran at the xhigh setting. This makes its performance even more impressive, as it achieved near-top scores without additional tuning. The league table from the final July 2026 Crucible shows gpt-5.6-sol leading at 95, with K3 close behind at 93.

This experiment underscores a critical insight for anyone considering AI for operational duties: the league is open, and choosing a model without your own rigorous testing is a gamble. The performance in real business scenarios can be quite different from chat demos or canned responses.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The Bottom Line: Do Your Due Diligence

AI models are evolving rapidly, but their true value lies in their ability to finish what they start, read your documents thoroughly, and stay honest under pressure. The recent live test proves that some frontier models are already capable of these traits—something you can verify yourself through live experiments. As the AI landscape shifts, the key for business leaders is to run your own wargames, not just rely on superficial demos. The league is open, and the best choice depends on how well an AI can handle your specific challenges, not just how well it chats.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Make Summer Dinners Easy with the Ninja Air Fryer 10QT DoubleStack XL

Discover the benefits of the Ninja Air Fryer 10QT DoubleStack XL for quick, family-sized summer meals. Perfect for grilling, roasting, and more!

Barcelona Rooftop Gardens Could Supply 31% Of City’s Tomatoes

A new study suggests that municipal rooftops in Barcelona could produce nearly a third of the tomatoes consumed locally, highlighting urban agriculture’s potential.

The Wisdom Of Wildflowers: Reflecting On Dolly Parton’s Timeless Song

Search interest in Dolly Parton’s song ‘Wildflowers’ is rising, reflecting its enduring appeal and cultural significance amid renewed public attention.

Master Summer Grilling with the Ninja Foodi Pro 5-in-1 Indoor Grill

Discover tips and hacks for getting perfect summer grilled meals with the Ninja Foodi Pro 5-in-1 Indoor Grill—your versatile indoor BBQ solution.