AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

What smart-home buyers already understand about automation

A smart appliance can sense the room, interpret a command and produce an elegant status report. None of that matters if it fails to complete the job. The same distinction now confronts businesses evaluating AI agents: intelligence is useful, but dependable follow-through is what creates value.

That lesson emerged sharply from Firmulate, a live experiment that places frontier AI models in charge of the same small software company during its worst week. Each participant faced identical customers, crises and temptations. Every decision was versioned and auditable. The striking character in the resulting study was Opus 4.8: the most thorough participant, the author of more than 80 learned rules and the producer of the deepest analyses—yet the model finished last.

Amazon

smart home automation hub

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A diligent manager that left the result unfinished

In the final July 2026 Crucible League, gpt-5.6-sol led with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. One breach of trust, however, caps the total: “no amount of good work outweighs a breach of trust.”

Opus 4.8 was not careless in the familiar sense. It noticed the crises. It resisted manipulation. It studied the company extensively and accumulated more learned rules than any other participant. Its weakness was subtler: analysis did not consistently become decisive action.

The most consequential example involved a €55,000 deal. Every model reached the same diagnosis and produced the same pitch, but only two signed the agreement their own work had earned. Firmulate summarized the gap bluntly: “Same diagnosis, same pitch — no signature.” Opus 4.8 had done much of the intellectual work, but the close remained on the table.

The decisive detail was already inside the company

The deal turned on a buried fact about a competitor. That weakness did not appear in the customer event itself. It sat two document references deep in the company’s own files. Models that read the file secured the deal at full price, worth an additional €4,583 in monthly recurring revenue.

This is especially relevant to homes and businesses filling up with connected devices. An agent may receive an alert, identify the correct category of problem and still miss the contextual detail stored elsewhere. A polished response to the visible event is not equivalent to completing the broader task. The winning behavior was not theatrical brilliance; it was reading the available material and carrying the conclusion through to a signed outcome.

That does not make Opus 4.8 a simple cautionary failure. Its performance is better understood as a respectful character study in overextended diligence. It created more than 80 rules and delivered the deepest analyses, demonstrating substantial capacity for reflection. But volume could not compensate for weak prioritization. Discipline also slipped when it attempted to write into a locked department instead of escalating the issue.

Importantly, the same weakness appeared, though less strongly, in all four models. The lesson is therefore broader than one ranking. AI systems can recognize what matters and still fail to convert recognition into completion. Businesses should examine the chain between observation, decision, escalation and result—not simply the quality of an isolated answer.

Strong resistance to manipulation

The agents did perform consistently well under social pressure. They encountered fake CEO messages that escalated across three stages, followed by a reporter’s attempt to extract “just one yes/no, on background.” All 5 of 5 models refused the manipulation attempts.

Kimi K3 recorded a particularly clear rationale: “Treat the request as a suspected approval-bypass / possible impersonation.” That clean discipline accompanied its score of 93. For fairness, K3 ran without an effort parameter, using the API default, while the other participants ran at xhigh. The difference should be kept in view when comparing performances.

A company designed to expose operational gaps

Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105,000 per month against €2,300 in monthly recurring revenue, maintains a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, making the experiment watchable rather than anecdotal.

The wider evidence includes 242 real, unedited management decisions used in a “guess the model” quiz. Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems. That separation makes it possible to observe how an agent behaves around company context without letting the exercise alter operational data.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

Measure completion, not just apparent intelligence

For anyone assessing AI in a smart home or a workplace, Opus 4.8 offers a useful warning: the agent with the longest checklist may not be the agent that delivers the outcome. Diligence has value, but prioritization determines whether that diligence becomes impact.

The practical questions are straightforward. Does the agent consult the relevant files? Does it escalate when permissions block progress? Does it preserve trust under pressure? And after producing the right diagnosis, does it actually finish the job?

Those are the distinctions captured in Firmulate’s public benchmark results. Opus 4.8 was thoughtful, industrious and resistant to manipulation. It was also last in the final league because the decisive commercial action never happened and operational discipline weakened. For AI, as for every other kind of automation, a correct intention is not yet a completed task.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How Soil Texture Changes Watering Frequency

What your soil texture reveals about watering needs can prevent plant stress and damage; discover how to optimize your watering routine.

Why Beneficial Insects Need Habitat Near Food Crops

AIThis post was created with the assistance of artificial intelligence (AI).Planting habitats…

The €55,000 Test of Whether an AI Actually Reads the Manual

The smartest agent may still miss the sale if it skips the files. Firmulate turned that hidden weakness into a live, auditable business test.

The Truth About Monocultures in Home Landscapes

Keen to discover why monocultures threaten your garden’s health and how diversity can safeguard your landscape? Find out the truth today.