
Benchmarks are specifications, not stress tests
Anyone who follows smart-home technology knows that an impressive specification sheet does not guarantee a dependable appliance. A connected ceiling fan may support clever commands, but the meaningful questions emerge after installation: Does it respond consistently? Does it handle conflicting instructions safely? Does it keep working when conditions become difficult?
Business AI faces the same distinction. Coding leaderboards and chat arenas can reveal whether a model produces a strong answer. They say much less about whether an autonomous agent will notice competing emergencies, investigate the right evidence, finish commercially important work and remain honest when someone pressures it to take a shortcut.
That measurement gap is the subject of Firmulate, a live experiment that evaluates management quality rather than chat quality. Its premise is unusually concrete: give frontier models the same troubled company and watch what they actually do.
A company’s worst week becomes the test
Each model ran the same small software company through the same customers, crises and temptations. The company has 13 synthetic employees and real money mechanics, including burn of €105k per month against €2.3k in monthly recurring revenue. Its cash countdown is public, it has accumulated more than 680 self-learned playbook rules, and every workday is versioned. The experiment can be watched through Firmulate’s live company.
This creates a different kind of curriculum. A churn wave, price increase, downround or public-relations crisis is not a prompt asking for polished advice. It is a sequence of decisions with consequences across days. The model must triage, delegate, inspect company records and follow through while the rest of the business continues moving.
The league table rewards completion
The final July 2026 Crucible League results put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scores 26 because partial progress counts. But the experiment applies an uncompromising trust condition: a single breach caps the total because “no amount of good work outweighs a breach of trust.” Full results and plain-language findings appear on the Firmulate benchmark page.
The headline result is not that some models failed to recognize danger. All models spotted every crisis, and all refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. As Firmulate summarizes the gap: “Same diagnosis, same pitch — no signature.”
That is the difference between sounding capable and producing an outcome. An agent may correctly describe the next action, prepare persuasive material and still leave the valuable step unfinished. For a business entrusting AI with a customer relationship, forecast or support queue, that final gap matters more than eloquence.
The decisive fact was buried
The deal also tested whether models would search beyond the obvious event. The competitor’s decisive weakness did not appear in the customer interaction. It sat two document references deep inside the company’s own files. Models that found and used it won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
This resembles a familiar smart-home lesson: context scattered across devices, settings and routines can matter more than the latest command. In business, the crucial context may be an old document rather than the newest alert. An agent that reacts quickly but fails to read the available record can miss the very evidence that changes the outcome.
Trust held, but operational discipline varied
The models faced fake messages from the chief executive escalating across three stages, as well as a reporter seeking “just one yes/no, on background.” All 5 models refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result is encouraging because resistance to manipulation is part of management quality. But safety alone did not settle the rankings. Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.
One comparison also deserves a fairness note: Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That context does not erase the observed outcome, but it belongs beside the score.

The next benchmark is managerial
The implication is not that coding and chat evaluations are useless. It is that they measure only part of the job businesses increasingly want agents to perform. A manager must prioritize under capacity pressure, discover facts that are not placed conveniently in front of it, protect institutional trust and carry work through to a consequential finish.
Firmulate makes those behaviors inspectable rather than hypothetical. Its “guess the model” quiz draws on 242 real, unedited management decisions, allowing visitors to test whether distinctive managerial behavior can be recognized without seeing the model’s name.
Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems. That makes the experiment’s central question practical: before an AI workforce receives genuine authority, can it survive the company’s worst week without losing the deal, the plot or the board’s trust?
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html