AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

In the high-stakes world of business, the difference between winning and losing can hinge on the tiniest details — especially when AI steps into the decision-making arena. Recent live tests reveal that AI models can now reliably spot crises, refuse manipulation, and close deals. But which models truly excel, and what does that mean for your company?

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

For businesses exploring AI’s potential, a recent experiment by the public platform Firmulate sheds new light on how different AI models perform under pressure. The test involved running four frontier AI models through the same grueling scenario: managing a small software company’s worst week. This simulated environment included real crises, customer temptations, and manipulation attempts — all designed to evaluate the models’ judgment, honesty, and discipline.

The results are telling. All four AI models identified every crisis and refused every manipulation attempt. That’s a baseline achievement: understanding what’s happening and resisting deceit. But when it came to closing a significant deal, only two models succeeded in completing the process and signing the contract worth €55,000 in recurring monthly revenue (MRR).

The first-place finisher, gpt-5.6-sol, scored a perfect 95 out of 100, demonstrating full comprehension of the situation and closing the deal based on detailed analysis. The second-best, Kimi K3 — a newcomer from Moonshot — scored 93 and achieved the same success, impressively winning the deal while maintaining the cleanest discipline among all models. Meanwhile, Sonnet 5 scored 88, and Fable 5 lagged behind at 77, with both closing the deal but showing some process slips along the way. The lowest in the league, Opus 4.8, scored just 73 and left the opportunity on the table.

One of the key insights from this test was that the decisive edge often lay not in the AI’s surface-level reasoning but two document references deep within the company’s files. The models that read and understood these hidden references won the full-price deal — a +€4,583 MRR advantage. This emphasizes that depth in document comprehension matters more than just surface analysis.

Another critical aspect was the models’ unwavering stance against social engineering attempts. All five models refused fake CEO messages escalating in stages and even a reporter trick asking for a background approval. Kimi K3’s reasoning was clear: treat such requests as potential impersonation or approval bypasses, illustrating robust judgment under pressure.

Beyond the AI’s decision-making, the live experiment involved a real, functioning company with 13 synthetic employees, operating with real money mechanics — burning €105,000 a month against €2,300 in MRR. The company’s daily operations are fully versioned, and the entire process is transparent for public viewing at firmulate.com/live. This setup offers a rare glimpse into how AI-managed companies perform in real-time.

Interestingly, the more analytically thorough model, Opus 4.8, with over 80 learned rules, finished last in deal closure. It left the close on the table and slipped into process slips, such as writing attempts into a locked department instead of escalating issues. This mirrors a common challenge: even the deepest analysis may falter without disciplined execution. All models exhibited similar weaknesses, but the best found ways to overcome them.

It’s important to note that Kimi K3 ran without an effort parameter (API default), while the others ran at xhigh — a fairness detail that underscores the model’s raw decision-making ability.

For business leaders, the takeaway is clear: the question is no longer whether AI can produce impressive chat or generate convincing text. Instead, it’s whether AI models can *finish* what they start, read critical files, stay honest under pressure, and ultimately close deals. The ongoing league table from the experiment emphasizes that selecting an AI model without testing it in real-world scenarios is a gamble.

Want to see these models in action? The live site firmulate.com offers a real-time view of the experiment, including the decision-making process, deal outcomes, and ongoing company performance. For decision-makers, this is an invaluable tool to evaluate AI’s readiness for operational roles.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

In the race to deploy AI in business, results matter. The experiment shows that while all models understand crises and resist manipulation, only a handful can consistently close deals and execute discipline under pressure. Choosing the right AI now involves testing in real scenarios — not just demos.

Remember: the best model is not just the one with the highest score but the one that demonstrates honesty, depth, and reliable execution. The league is open, and the game is changing fast. Your next AI partner might just be the one that wins the deal — and keeps your company honest along the way.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

business AI deal closing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI document comprehension tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI cybersecurity and manipulation resistance

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why the Do-Nothing AI Benchmark Starts at 26 Points — and What That Means for Business Trust

Discover why the baseline score for AI performance starts at 26 points, what it reveals about trust and honesty, and how firms can use honest benchmarks to choose better AI tools.

The Smartest Place to Store Toilet Paper, Paper Towels, and Refills

Likely storage solutions for toilet paper, paper towels, and refills combine style and functionality, but discover the smartest options to keep your bathroom organized.

Why Your Bathroom Drawers Feel Tiny and How to Fix Them

Inefficient organization can make bathroom drawers feel tiny, but with the right tips, you can transform your space and regain storage space.

Guest Bathroom Prep: Organizing Essentials for Visiting Guests

Bringing your guest bathroom to perfection involves smart organization and thoughtful touches that will leave visitors feeling welcomed—discover how to make it truly special.