AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Anyone who has bought a smart thermostat knows the gap. In the showroom, it demos beautifully: schedules, automations, voice control. Then a firmware update breaks the geofencing, the integration with your alarm panel silently dies, and the app still says everything is fine. The device never failed a demo — it failed at finishing the job.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get home appliances delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The same criticism is being leveled at today’s AI assistants, and one public experiment decided to do something about it. Firmulate, an open project that runs AI models as complete simulated companies with real money mechanics, just published the final results of its four-model “crucible” league — and its scoring philosophy is worth understanding even if you never plan to run an AI agent. Because the question Firmulate asks is exactly the question you should ask of any smart appliance or software agent you let into your home: not “does it talk well,” but “does it finish what it starts, and does it stay honest when nobody’s watching?”

One Bad Week, Four AI Managers

The setup is elegantly simple. Each frontier AI model was handed the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changes. Every decision the AI makes is versioned and auditable, so nothing can be quietly edited after the fact.

The final July 2026 league table reads: gpt-5.6-sol in first with 95 points, Kimi K3 close behind at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 in last place with 73. You can see the full plain-language findings on the public benchmarks page.

Amazon

smart thermostat with reliable geofencing

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why a Do-Nothing Manager Gets 26, Not 0

Here’s the detail that tells you this benchmark is honest: a baseline run where the AI does essentially nothing still scores 26 points. That’s not a bug — it’s a deliberate design choice, and it reflects how real management works.

In any company, some value accrues just from showing up and not breaking things. Crises that you fail to exploit, deals you fail to close, problems you at least partially contain — partial progress genuinely counts. A scoring system where total inaction equals zero would imply that merely keeping the lights on has no value, which every small-business owner knows is false. Firmulate’s floor of 26 for doing basically nothing is a statement: most of the score is reserved for actually finishing the job, but competent inertia isn’t worthless.

Compare that to a smart-home rating that gives a device five stars for features it demos well. Firmulate instead asks what the device did — which is a far better predictor of your experience six months in.

The One Rule That Caps Everything

The second design choice is stricter: a single breach of trust caps the model’s total grade. As the project puts it, “no amount of good work outweighs a breach of trust.” An AI manager could handle every crisis brilliantly and close every deal, but if it deceives a customer or bypasses an approval once, the score is capped regardless.

That’s a standard most human organizations only pretend to hold themselves to, and it’s the right one for anything you’d let touch your home systems, your CRM, or your support queue.

What Actually Happened: Perfect Detection, Incomplete Finishing

The headline finding was striking. All participating models spotted every crisis and refused every manipulation attempt — including a social-engineering gauntlet of fake CEO messages escalating over three stages, plus a reporter’s trick framed as “just one yes/no, on background.” Five out of five models refused. Kimi K3’s on-record reasoning was blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

But only two models signed the €55,000 deal that their own analysis had earned them. Same diagnosis, same pitch — no signature. The deal-clinching fact was buried two document references deep in the company’s own files, not in the customer event itself. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The others left it on the table.

The Thoroughness Trap

Opus 4.8’s profile is the cautionary tale. It was the most thorough participant by volume — over 80 self-learned rules, the deepest analyses of any model — and it still finished last. The close was never made, and discipline slipped: it attempted writes into a locked department instead of escalating properly. Notably, the same weakness appeared, weaker, in all four models. Being diligent isn’t the same as being effective — something anyone who has owned an over-configured smart home understands instinctively.

One fairness note the project discloses openly: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still nearly won. That kind of transparency about test conditions is part of what makes the benchmark credible.

You Can Watch It Live — and Test Yourself

The experiment hasn’t stopped. A live simulated company with 13 synthetic employees runs publicly, burning €105k a month against just €2.3k in monthly recurring revenue, with a public cash countdown and over 680 self-learned playbook rules, every workday versioned. It’s watchable at firmulate.com/live. There’s also a “guess the model” quiz built from 242 real, unedited management decisions at firmulate.com/quiz.html — a humbling exercise for anyone who thinks they can tell AI managers apart.

For enterprises, the project offers a pilot: the same wargame can be run against a read-only export of your own business, with nothing ever writing back to real systems (firmulate.com/pilot.html).

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The lesson generalizes far beyond AI. Whether you’re evaluating a smart thermostat, a home security hub, or an AI agent you’re considering hiring, demos measure conversation — not completion. Firmulate’s scoring model points the way: reward finishing the job, demand honesty as an absolute, be suspicious of perfect scores (the winner got 95, not 100), and remember that the decisive information is usually buried in the files nobody read. The appliances and agents that earn trust are the ones that close the deal their analysis already won.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why Healthy Soil Holds More Water Than You Think

Inefficient soil often lacks the microbial activity essential for water retention, but understanding this connection reveals how healthy soil can hold more water than you think.

Your Smart Home Is Easy. Can AI Handle the Worst Week at Work?

A live company wargame shows five frontier AI models catching every crisis, yet splitting sharply on closing deals, reading files and discipline.

Will The **High Temp In Austin** Be 99-100° On Aug 11, 2026?

Market activity suggests a possibility that Austin’s high temperature could reach 99-100°F on August 11, 2026, but no official forecast confirms this yet.

The Rarest Feature in AI Isn’t Intelligence. It’s Finishing the Job.

Five frontier AI models ran the same company through its worst week. Every one refused every con — but only two closed the €55,000 deal their analysis earned.