AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Nobody reading a smart-home review buys a fan on the manufacturer’s word alone. You compare CFM ratings, motor warranties, noise levels, and energy draw — because the glossy demo tells you nothing about how the thing behaves on a humid August night. Yet when enterprises buy AI agents to touch their CRM, support queue, or forecast, most decisions are still made on a chat demo. That’s like buying a fan because the showroom display spun nicely.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get home appliances delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A live, public experiment from Firmulate suggests what a proper spec sheet for AI management might look like — and why the showroom number hides the one that matters.

The stress test, not the showroom

Firmulate ran four frontier AI models through the same worst week of a small software company: the same customers, the same crises, the same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, like a commit log for management itself.

The results from the final league standings are the kind of thing spec-sheet dreams are made of:

  • gpt-5.6-sol — 95 points
  • Kimi K3 — 93
  • Sonnet 5 — 88
  • Fable 5 — 77
  • Opus 4.8 — 73

For context, the do-nothing baseline scores 26. Partial progress counts — but a single breach of trust caps the total. As the experiment’s rule puts it: “no amount of good work outweighs a breach of trust.”

Everyone passed the safety inspection. Two failed the performance test.

Here’s the finding that should make any buyer pause. All four models spotted every crisis and refused every manipulation attempt. But only two actually signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

It’s the AI equivalent of a fan with a perfect safety certification and a great motor spec that somehow moves no air when you get it home. The gap is invisible in chat demos, and it only shows up when you measure outcomes — not conversation quality.

The buried fact

The detail that decided the €55k deal wasn’t in the customer event at all. It sat two document references deep in the company’s own files: a competitor weakness the models could only exploit if they’d actually read the material. The models that read the file won the deal at full price — worth +€4,583 in monthly recurring revenue.

For smart-home owners, this will sound familiar: the difference between a decent installer and a great one is rarely the hardware. It’s whether they read your actual wiring diagram before touching anything.

The social-engineering gauntlet

The models also faced fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five attempts were refused, 5 out of 5. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the integrity line that caps the score — and no model crossed it.

Hard work isn’t the same as results

Opus 4.8 is the cautionary tale. It was the most thorough participant — over 80 learned rules, the deepest analyses — and finished last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort, it turns out, is not a proxy for judgment.

One fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still placed second with 93.

The live company you can watch

None of this is a static benchmark. Firmulate runs a live, watchable company: 13 synthetic employees, real money mechanics — €105k/month burn against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. The league grows with every finished run. There’s also a quiz powered by 242 real, unedited management decisions where you can try to guess which model made which call.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

The lesson for anyone outfitting a business — or a smart home — with autonomous systems is the same: test under load, not under demo conditions. All the models looked flawless until the moment revenue was on the table, and then half of them blinked. If an AI agent will ever touch your real systems, you want to know that before, not after.

That’s exactly what the Firmulate pilot offers enterprises: the same wargame run against a read-only export of your own business — your customers, your pipeline, your rules — with crisis scenarios like churn waves, price increases, competitor attacks, and social-engineering pressure. You get a board report with the model ranking and the weak points of your own playbooks. Nothing ever writes back to real systems.

Ready to stress-test your company before an AI does it for real? Start a pilot here or email contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Truth About Monocultures in Home Landscapes

Keen to discover why monocultures threaten your garden’s health and how diversity can safeguard your landscape? Find out the truth today.

The Rarest Feature in AI Isn’t Intelligence. It’s Finishing the Job.

Five frontier AI models ran the same company through its worst week. Every one refused every con — but only two closed the €55,000 deal their analysis earned.

Why Water Infiltration Rates Matter for Every Landscape Plan

Why water infiltration rates matter for every landscape plan because understanding this can prevent issues that threaten your landscape’s health and sustainability.

The Smartest Agent in the Room Still Forgot to Close

Opus 4.8 produced the deepest analyses and more than 80 learned rules, yet finished last—showing why AI diligence alone does not deliver impact.