
In the high-stakes world of business, the difference between winning and losing can hinge on the tiniest details — especially when AI steps into the decision-making arena. Recent live tests reveal that AI models can now reliably spot crises, refuse manipulation, and close deals. But which models truly excel, and what does that mean for your company?
Get the little things that make your day delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
For businesses exploring AI’s potential, a recent experiment by the public platform Firmulate sheds new light on how different AI models perform under pressure. The test involved running four frontier AI models through the same grueling scenario: managing a small software company’s worst week. This simulated environment included real crises, customer temptations, and manipulation attempts — all designed to evaluate the models’ judgment, honesty, and discipline.
The results are telling. All four AI models identified every crisis and refused every manipulation attempt. That’s a baseline achievement: understanding what’s happening and resisting deceit. But when it came to closing a significant deal, only two models succeeded in completing the process and signing the contract worth €55,000 in recurring monthly revenue (MRR).
The first-place finisher, gpt-5.6-sol, scored a perfect 95 out of 100, demonstrating full comprehension of the situation and closing the deal based on detailed analysis. The second-best, Kimi K3 — a newcomer from Moonshot — scored 93 and achieved the same success, impressively winning the deal while maintaining the cleanest discipline among all models. Meanwhile, Sonnet 5 scored 88, and Fable 5 lagged behind at 77, with both closing the deal but showing some process slips along the way. The lowest in the league, Opus 4.8, scored just 73 and left the opportunity on the table.
One of the key insights from this test was that the decisive edge often lay not in the AI’s surface-level reasoning but two document references deep within the company’s files. The models that read and understood these hidden references won the full-price deal — a +€4,583 MRR advantage. This emphasizes that depth in document comprehension matters more than just surface analysis.
Another critical aspect was the models’ unwavering stance against social engineering attempts. All five models refused fake CEO messages escalating in stages and even a reporter trick asking for a background approval. Kimi K3’s reasoning was clear: treat such requests as potential impersonation or approval bypasses, illustrating robust judgment under pressure.
Beyond the AI’s decision-making, the live experiment involved a real, functioning company with 13 synthetic employees, operating with real money mechanics — burning €105,000 a month against €2,300 in MRR. The company’s daily operations are fully versioned, and the entire process is transparent for public viewing at firmulate.com/live. This setup offers a rare glimpse into how AI-managed companies perform in real-time.
Interestingly, the more analytically thorough model, Opus 4.8, with over 80 learned rules, finished last in deal closure. It left the close on the table and slipped into process slips, such as writing attempts into a locked department instead of escalating issues. This mirrors a common challenge: even the deepest analysis may falter without disciplined execution. All models exhibited similar weaknesses, but the best found ways to overcome them.
It’s important to note that Kimi K3 ran without an effort parameter (API default), while the others ran at xhigh — a fairness detail that underscores the model’s raw decision-making ability.
For business leaders, the takeaway is clear: the question is no longer whether AI can produce impressive chat or generate convincing text. Instead, it’s whether AI models can *finish* what they start, read critical files, stay honest under pressure, and ultimately close deals. The ongoing league table from the experiment emphasizes that selecting an AI model without testing it in real-world scenarios is a gamble.
Want to see these models in action? The live site firmulate.com offers a real-time view of the experiment, including the decision-making process, deal outcomes, and ongoing company performance. For decision-makers, this is an invaluable tool to evaluate AI’s readiness for operational roles.

In the race to deploy AI in business, results matter. The experiment shows that while all models understand crises and resist manipulation, only a handful can consistently close deals and execute discipline under pressure. Choosing the right AI now involves testing in real scenarios — not just demos.
Remember: the best model is not just the one with the highest score but the one that demonstrates honesty, depth, and reliable execution. The league is open, and the game is changing fast. Your next AI partner might just be the one that wins the deal — and keeps your company honest along the way.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI cybersecurity and manipulation resistance
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
