
In the world of AI, the idea of a perfect score is tempting — but in real-world testing, even the most straightforward tasks reveal honesty as the ultimate currency. Imagine a scenario where an AI, given a simple ‘do-nothing’ baseline, scores 26 out of 100. That might sound underwhelming, but it’s actually a sign of a healthy, trustworthy benchmark. For business leaders, understanding this nuance can be the difference between deploying AI that truly helps and one that merely looks good on paper.
Get your wardrobe favorites delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Reality of AI Benchmarks: More Than Just Scores
When AI models are measured, it’s tempting to think of the scores as a simple ladder — higher is better, right? But the recent experiment conducted by Firmulate paints a different picture. In a detailed test called the Crucible League, four frontier models were put through the same tough week at a mock software company. The goal: see if they could handle crises, avoid manipulation, and make honest decisions.
Remarkably, all four models identified every crisis and refused every attempt at manipulation, such as fake CEO messages or bribery. Yet, only two signed a €55,000 deal based on their own analysis. The other two, despite doing the same diagnosis and pitch, left the deal on the table. The scores reflected this: the best model scored 95, the weakest scored 77. But crucially, the baseline — a do-nothing approach that simply did nothing — scored 26.
As an affiliate, we earn on qualifying purchases.
Why Does a Do-Nothing Baseline Score 26?
At first glance, it might seem odd that a model doing nothing scores above zero. The reason lies in the experiment’s methodology. Even the simplest benchmark run involves some partial progress — for example, reading documents or making decisions that don’t necessarily lead to action. These partial steps count toward the final score, meaning that the baseline isn’t just a null state but a reflection of minimal engagement.
Furthermore, the experiment caps the total score if the model breaches trust, meaning that a single dishonest act — even if the model performs well otherwise — can set a ceiling. This ensures that honesty isn’t just an added bonus but a core requirement for a high score. It’s a way to measure not just what the AI can do, but whether it can be trusted to do the right thing.
trustworthy AI model evaluation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Does This Mean for Business AI Deployment?
For companies considering AI systems for customer service, internal decision-making, or automation, this experiment is a wake-up call. The key isn’t just that an AI can produce convincing outputs — it’s whether it can finish the job honestly and reliably. The models in the experiment saw every crisis, refused manipulative pleas, and only signed deals when they were confident in their analysis.
This emphasis on honesty and process integrity is vital because it directly affects the bottom line. In the live test, a single document reference hidden deep inside the company’s files was the decisive factor in closing a deal at full price — a detail no chatbot demo would reveal. That’s the kind of thoroughness that can make or break a real-world business decision.
business AI decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Human Element: Trust and Discipline
Interestingly, the experiment also surfaced human-like weaknesses. In the case of OPUS 4.8, despite being the most thorough participant with over 80 learned rules, it left the deal unclosed and slipped in discipline — writing attempts into a locked department rather than escalating. The takeaway? Even sophisticated AI can falter if not properly managed and disciplined.
In practice, this means that companies should not rely solely on performance scores but also consider how AI models behave under pressure. The experiment’s fairness note — that models run at different effort levels — highlights that consistency and discipline are crucial. A model that scores high only at high effort isn’t necessarily the best choice for day-to-day operations.
AI honesty and integrity testing
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Bottom Line: Honest Benchmarks Lead to Better Business Outcomes
This entire experiment underscores an important truth: the value of AI in business isn’t just about surface-level capabilities. It’s about whether AI systems can genuinely trust, read deeply, and make honest decisions under pressure. The fact that the baseline scores 26 points — despite doing nothing — serves as a reminder that even simple, honest behavior is a baseline standard for trustworthy AI.
For business leaders and decision-makers, the takeaway is clear: when evaluating AI, look beyond shiny demos and high scores. Focus on how models handle crises, resist manipulation, and whether they finish what they start. That’s the real measure of an AI’s readiness to serve your company in the long run.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
