
Imagine you’re hiring an assistant. You want someone honest, reliable, and able to handle crises without cutting corners. But what if your benchmark for trust isn’t zero, but 26 points out of 100? That’s the reality of the latest AI performance testing — and it reveals much about how AI is shaping business integrity today.
Get the little things that make your day delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Truth Behind the Baseline Score
In recent experiments conducted by Firmulate, four advanced AI models were put through a simulated week in the life of a small software company. The goal wasn’t just to see if they could chat well — it was about managing crises, making decisions, and maintaining honesty under pressure. The results? All four models identified every crisis and refused every manipulation attempt. Yet, surprisingly, their scores started at 26 points even before doing any work.
Partial Progress Counts — and Why That Adds Up
You might wonder: why doesn’t the baseline score start at zero? The reason is that even doing nothing in this test environment yields partial credit — 26 points, to be exact. This partial credit recognizes that the AI models are not blank slates; they possess inherent capabilities. They can recognize crises, refuse manipulation, and read documents. These are baseline skills that matter in real-world scenarios, where an AI’s ability to even understand the context is crucial.
Trust Breaches Limit the Total Score
However, there’s a catch — if an AI model breaches trust even once during the test, it cannot recover. The experiment’s rules cap the score, emphasizing that no amount of good performance can outweigh a single act of dishonesty. This mirrors real business expectations: trust, once lost, can’t be fully regained through partial wins.
AI performance benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Experiment Revealed About AI in Business
Each AI model was tested against the same scenario: managing a small company with real money, real crises, and real temptations. The models successfully identified all crises and refused manipulative tactics, including social engineering attempts like fake CEO messages and reporter tricks. Their discipline in refusing manipulative requests was unanimous — five out of five models refused to be tricked.
The Hidden Weakness: Reading Deeper Files
The decisive factor wasn’t just in quick responses or surface-level decisions. The models that read two documents deep into the company’s records managed to secure the €55,000 deal at full price, adding €4,583 in monthly recurring revenue. Those that didn’t read that far left the deal on the table, missing out on substantial business value.
Discipline Under Pressure: The Opus 4.8 Profile
The most meticulous participant, Opus 4.8, demonstrated that thoroughness matters — but it wasn’t enough to close the deal. Despite analyzing over 80 rules and conducting deep assessments, its discipline slipped, and it left the close opportunity on the table. The same weakness appeared, albeit weaker, in all models tested, highlighting a common challenge in AI decision-making under complex scenarios.
As an affiliate, we earn on qualifying purchases.
Implications for Business Trust and AI Adoption
This experiment underscores a vital reality: AI models are not just about generating convincing chat responses. Their true value lies in their ability to stay honest, read deeply, and follow through. For businesses considering AI for critical functions — from customer management to decision-making — the question isn’t just whether the AI can talk well, but whether it can truly finish what it starts.
The Role of Transparent Benchmarks
Firmulate’s live experiment offers a transparent, auditable way to evaluate AI performance in real business contexts. The benchmarks reveal the exact points where models excel or stumble — and why trust is a hard-earned commodity in AI systems. The scores, like the 26 baseline points, serve as honesty indicators, not just performance metrics.
As an affiliate, we earn on qualifying purchases.
What This Means for Your Business
If you’re considering deploying AI in your company, remember: a high chat score doesn’t guarantee trustworthy behavior. Look for models that can handle crises, recognize deeper documents, and refuse manipulation — and that can do so consistently. The value of an AI isn’t just in its words, but in its integrity and discipline when it matters most.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
