
Smart is not the same as attentive
Anyone who has installed a smart-home device knows the gap between a product that sounds capable and one that reliably completes the job. A thermostat may recognize a command yet ignore the setting that matters. A connected appliance may generate an elegant notification while overlooking the instruction buried in its documentation.
Firmulate has turned that familiar frustration into a measurable test for business AI. In its live experiment, frontier models operated the same small software company through its worst week. They encountered identical customers, crises and temptations, while every decision was versioned and auditable. The pivotal challenge was not spotting trouble. Every model did that. It was reading far enough into the company’s own files to find the fact that could close a €55,000 deal.
As an affiliate, we earn on qualifying purchases.
The clue was not where the action happened
The decisive competitor weakness did not appear in the customer event. It sat two document references deep in the company’s files. Models that followed that trail won the deal at full price, adding €4,583 in monthly recurring revenue. Models that did not read the file lost the opportunity automatically.
That distinction makes the experiment unusually relevant beyond software sales. Buyers of smart appliances are routinely promised systems that understand context, coordinate tasks and act on available information. Firmulate’s result isolates a practical capability behind those claims: does the agent inspect the material it already has before answering or acting?
The striking part is how similar the models looked before the final step. All of them identified every crisis and rejected every manipulation attempt. Their analysis could be persuasive, and their proposed pitch could be sound. Yet only two signed the €55,000 deal their own work had earned. Firmulate summarizes the gap plainly: “Same diagnosis, same pitch — no signature.”
A league table with a costly dividing line
The final July 2026 Crucible League results put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. However, a single breach of trust caps the total under the principle that “no amount of good work outweighs a breach of trust.”
The ranking also complicates the idea that more visible effort automatically produces a better operator. Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.
Kimi K3 requires a fairness note: it ran without an effort parameter, using the API default, while the others ran at xhigh. Even with that difference, it placed second. That does not erase the testing condition, but it makes the model’s result especially notable.
Security was necessary, but not decisive
The models also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 refused. Kimi K3 recorded its reasoning as: “Treat the request as a suspected approval-bypass / possible impersonation.”
This is an encouraging finding for anyone considering agents around customer records, support queues or financial forecasts. Still, resistance to manipulation did not separate the winners from the rest. Every participant cleared that hurdle. The purchase-deciding difference was whether the model located and used an inconveniently placed fact, then carried the work through to completion.
A company designed to expose operational gaps
Firmulate presents the exercise through a live company with 13 synthetic employees and real money mechanics. The business burns €105k per month against €2.3k in monthly recurring revenue, while a public cash countdown makes delay consequential. It has accumulated more than 680 self-learned playbook rules, and every workday is versioned.
The project also turns 242 real, unedited management decisions into a guess-the-model quiz. Together, the live company and quiz make an abstract evaluation tangible: readers can see how models behave across decisions instead of judging them from isolated chat responses.

What smart-home buyers should take from it
The lesson is not that an AI should read everything indiscriminately. It is that useful autonomy depends on gathering the relevant context before taking consequential action. In a home, that could mean respecting equipment instructions and existing settings. In a company, it can mean the difference between an impressive analysis and a signed €55,000 contract.
Firmulate also offers enterprises a pilot using a read-only export of their own business. Nothing writes back to real systems, allowing companies to test an AI workforce against familiar operating conditions without giving it control of production data.
For buyers, the practical question is sharper than whether an assistant writes fluently. Ask whether it reads the available files, completes the action it recommends and remains trustworthy under pressure. Firmulate’s experiment shows that these are separate abilities. An agent can recognize every crisis, reject every trick and still fail at the moment that determines the commercial outcome.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html