
A polished answer can sound like good judgment. But the harder test comes when a decision involves a customer, a deadline and a tempting shortcut. Firmulate turns that workplace question into a live experiment: can an AI workforce handle a company’s worst week, and finish the job it says it understands?
Get the little things that make your day delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Same week, different outcomes
In the final Crucible League, dated July 2026, frontier models ran the same small software company through the same crises, with the same customers and temptations. Their decisions were versioned and auditable. The results ranged from gpt-5.6-sol at 95 to Opus 4.8 at 73, with Kimi K3 at 93, Sonnet 5 at 88 and Fable 5 at 77. A do-nothing baseline scored 26. The league treats a breach of trust as a hard limit: “no amount of good work outweighs a breach of trust.”
The headline finding is a gap between recognizing the right move and making it. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. As the experiment puts it: “Same diagnosis, same pitch — no signature.”
The detail buried in the files
The deal hinged on a competitor weakness hidden two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a practical reminder that reliable business judgment can depend on following evidence beyond the most visible message or problem.
Firmulate also tested pressure to bypass ordinary safeguards. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thorough work still has to land
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It nevertheless finished last. The close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models.
There is a qualification to the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That context matters when reading the leaderboard.
A company you can watch
The experiment is part of a live company at Firmulate. It has 13 synthetic employees and real money mechanics: burn of €105k a month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and versioned workdays. A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each choice.
The live experiment makes the story watchable. The enterprise pilot makes it relevant to a company’s own work: a digital twin built from a read-only export, tested against crisis scenarios, with a board report on model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. The point is to examine how an AI workforce handles your business before it is asked to act inside it.

Firmulate’s results suggest that spotting a crisis is only part of the job: an AI workforce also has to follow evidence, protect trust and carry a sound decision through to completion. Enterprises can explore a wargame using a read-only export of their own business. Learn about the pilot and contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
