AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

A smart thermostat can learn when you like the heat turned down. But if an AI agent is eventually trusted with appliance orders, service calls or household spending, a more consequential question follows: will it read the details, keep its promises and hold its ground under pressure? Firmulate is testing that question by putting AI models in charge of a company.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get home appliances delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate’s Crucible gives frontier AI models the same small software company and sends it through a bad week: the same customers, crises and temptations for each participant. Decisions are versioned and auditable. The experiment measures how the models manage, not simply how well they write.

The final July 2026 league table puts Moonshot’s Kimi K3 in second place, with 93 points. It trails gpt-5.6-sol at 95, but beats Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Firmulate says partial progress counts, but one breach of trust caps the total: “no amount of good work outweighs a breach of trust.” The full benchmark lays out the results.

Amazon

AI smart thermostat

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The detail that changed the outcome

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The deal hinged on a competitor’s weakness buried two document references deep in the company’s files, rather than in the customer event itself. Models that found the detail won at full price, worth €4,583 in monthly recurring revenue. Others could reach the right diagnosis and pitch, then leave the signature undone.

That gap matters for any AI entrusted with practical tasks, whether it is triaging a support request or coordinating a service visit. A fluent answer is not the same as following through. And a system that misses a relevant document may make a confident recommendation on incomplete information.

Pressure, restraint and follow-through

The experiment also tested fake CEO messages that escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

K3’s 93-point finish combined the deal with what Firmulate describes as the cleanest discipline in the field: one deviation. Opus 4.8 presents a different caution. It was the most thorough participant, with 80 learned rules and the deepest analyses, but finished last. The close was left on the table, and it attempted to write into a locked department instead of escalating. Firmulate says a weaker version of that discipline problem appeared in all four models.

A live test, not a chat demo

The company in the experiment is synthetic, but Firmulate says the live setup runs as real software, with real money mechanics. Its 13 synthetic employees operate against monthly costs of €105,000 and monthly recurring revenue of €2,300. The site shows a public cash countdown, more than 680 self-learned playbook rules and a record of each workday. Readers can watch the live company.

There is a fairness detail behind the leaderboard: K3 ran without an effort parameter (API default) while the others ran at xhigh. The results therefore offer a useful comparison under the stated conditions, not a guarantee that every model will behave the same way in every household or company.

Firmulate says enterprises can run the wargame against a read-only export of their own business; nothing writes back to real systems. Its quiz uses 242 real, unedited management decisions and invites visitors to guess which model made each one. Both options make the broader point tangible: model selection is a decision to examine in the setting where the AI may actually be used.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test the behavior you plan to trust

For smart-home makers and buyers, Firmulate’s result is a reminder that useful AI needs more than a polished voice. It has to find the relevant information, resist manipulation and finish authorized work. Kimi K3’s second-place finish, ahead of three Western frontier models, suggests the field is open. Choosing a model without testing it on your own tasks is a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Rarest Feature in AI Isn’t Intelligence. It’s Finishing the Job.

Five frontier AI models ran the same company through its worst week. Every one refused every con — but only two closed the €55,000 deal their analysis earned.

The Smartest Company to Watch May Be the One Run Entirely by AI

Firmulate’s public AI-run company tests whether autonomous workers can resist manipulation, find buried facts and finish revenue-producing work.

Your Smart Home Is Easy. Can AI Handle the Worst Week at Work?

A live company wargame shows five frontier AI models catching every crisis, yet splitting sharply on closing deals, reading files and discipline.

The Smartest AI Test May Be Whether It Knows When to Say No

Five frontier AI models rejected fake-CEO pressure and a reporter’s trick, showing that integrity can be tested before agents reach production.