AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get pool and patio gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

What if your next AI assistant could run your business — and actually finish what it starts?

Imagine deploying an AI that’s not just good at chatting but can handle your company’s toughest crises, read your documents thoroughly, and resist shady manipulations. That’s the promise—and the challenge—facing AI developers today. Recent live experiments reveal that some models are starting to deliver on this potential, outperforming others in real-world tests that mirror the chaos of actual business operations.

Amazon

AI business management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Live Experiment: Putting AI to the Test in a Real Business Setting

In a groundbreaking live experiment conducted by Firmulate, four frontier AI models were tasked with running a small software company through its most challenging week. This wasn’t a simulation in a chat interface; it was a real-time, auditable test that involved actual crises, customer decisions, and the temptation to cheat. The goal: see which AI could truly manage the company’s worst week — not just talk about it.

All four models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5—demonstrated a strong ability to recognize crises and refused every manipulation attempt, including social engineering tactics like fake CEO messages. Yet, the differences lay in their ability to complete critical tasks and close deals. Kimi K3, a newcomer from Moonshot, scored a 93 out of 100, just behind the top scorer, gpt-5.6-sol, with 95. The standout for K3 was its ability to find a buried security fact deep within the company’s files—a crucial detail that clinched a €55,000 deal, translating into +€4,583 monthly recurring revenue (MRR).

Cracking the Hidden Code: Kimi K3’s Edge

The experiment uncovered a buried weakness common among models: the failure to read and interpret key documents. While the top models identified critical facts buried two document references deep, others left the deal on the table or slipped discipline, showing weaker performance. K3’s success highlights an important insight: reading comprehension at depth can be decisive in real business outcomes, not just in chat quality.

Honesty and Integrity Under Pressure

The experiment also tested models against social engineering—fake CEO messages escalating in three stages, plus a reporter’s trick question. All models refused to approve or sign any fake requests, citing suspicion and adherence to protocol. Kimi K3’s on-record reasoning was clear: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates an encouraging level of honesty under pressure, a crucial trait for AI in business applications.

The Real Business Environment

The live setting involved 13 synthetic employees, real cash mechanics, and a public focus on daily performance metrics. The company burned €105,000 monthly against only €2,300 in MRR, with a public cash countdown actively ticking down. Every decision was versioned and auditable, and the entire process is observable at firmulate.com/live. This transparency underscores how AI models perform outside of scripted demos, facing genuine operational dilemmas.

The Lessons for Business Leaders

While the most thorough participant, Opus 4.8, scored the lowest (73), it shed light on discipline weaknesses, such as leaving deals on the table or failing to escalate issues. The results emphasize a vital point: the best AI models are those that not only recognize crises but also follow through consistently and honestly, especially under pressure.

Fairness Note and the Open League

It’s important to note that Kimi K3 ran without an effort parameter (the API default), while the other models ran at the xhigh setting. This makes its performance even more impressive, as it achieved near-top scores without additional tuning. The league table from the final July 2026 Crucible shows gpt-5.6-sol leading at 95, with K3 close behind at 93.

This experiment underscores a critical insight for anyone considering AI for operational duties: the league is open, and choosing a model without your own rigorous testing is a gamble. The performance in real business scenarios can be quite different from chat demos or canned responses.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The Bottom Line: Do Your Due Diligence

AI models are evolving rapidly, but their true value lies in their ability to finish what they start, read your documents thoroughly, and stay honest under pressure. The recent live test proves that some frontier models are already capable of these traits—something you can verify yourself through live experiments. As the AI landscape shifts, the key for business leaders is to run your own wargames, not just rely on superficial demos. The league is open, and the best choice depends on how well an AI can handle your specific challenges, not just how well it chats.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Is It Too Early To Think About Fall Decor? Because Joanna Gaines Just Dropped This Faux Foliage At Target, And I’m Obsessed Already

Joanna Gaines has introduced a new fall decor line at Target, featuring faux foliage and seasonal accents, sparking early excitement among shoppers.

Plant Matchmaker: This Romantic Prairie Pairing Bridges Summer And Fall In Shimmering Waves Of Color

A new garden pairing blends summer and fall blooms, creating a seamless transition of colors in prairie landscapes, according to horticultural experts.

Create Delicious Summer Treats with the Ninja NC301 CREAMi Ice Cream Maker

Learn how to make refreshing summer ice creams, gelatos, and smoothies easily with the Ninja NC301 CREAMi Ice Cream Maker in this step-by-step guide.

Ninja Foodi 10 Quart DualZone Air Fryer: Summer Family Feast Ready

Discover the Ninja Foodi XL Air Fryer — perfect for summer family meals with versatile cooking, dual-zone tech, and easy cleanup. A summer must-have!