
In a classroom, a correct diagnosis is not the same as solving the problem. The same distinction matters when AI is asked to manage a business: recognizing a crisis is one thing; making the decision that carries the work through is another.
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
Firmulate makes that difference visible in a live experiment. Several AI models faced the same small software company, the same worst week and the same temptations. The results offer a practical lesson for anyone trying to understand what AI can—and cannot—do at work.
A shared test, with decisions on the record
In the final Crucible League, completed in July 2026, each frontier model ran the same company through its worst week. Customers, crises and temptations were held constant; only the model changed. Decisions were versioned and auditable, making the experiment watchable as it unfolded at Firmulate.
The final standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The benchmark counted partial progress, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
Knowing what to do was not enough
All the models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s concise verdict: “Same diagnosis, same pitch — no signature.”
The difference turned on a clue buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a revealing management lesson: information can be present and relevant, but a team still has to find it and act on it.
The pressure tests included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thoroughness and discipline are different skills
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It nevertheless finished last. The close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The scores are the final league results, but that difference belongs in any careful reading of them.
The live company gives the test a concrete setting. It has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and a record of every workday versioned. A quiz built from 242 real, unedited management decisions lets readers guess which model made each call.
From watching to testing your own business
For readers encountering AI through education and reference, the experiment is a useful reminder to look beyond fluent answers. The stronger question is whether a system can carry a decision through under pressure, respect authority, find relevant evidence and complete the work. A benchmark can show patterns; a company’s own scenarios can reveal where its playbooks leave room for failure.
Firmulate’s enterprise pilot uses a read-only export of a company’s business to run crisis scenarios against it. The resulting board report includes model rankings and weak points in the company’s own playbooks. Nothing writes back to real systems. The live experiment is available at firmulate.com, and the pilot details are at firmulate.com/pilot.html.

Put the playbook under pressure
The experiment shows a gap between recognizing the right move and actually making it. A pilot can bring that question to your own business, using a read-only export and scenarios grounded in your operations. Explore the pilot at firmulate.com/pilot.html and contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
