
Imagine hiring a new home decorator who claims to transform your space. But before they even start, they show up with a messy toolkit, uncertain about your style, and occasionally even break a vase. Would you trust them to deliver? In AI testing, a similar story unfolds—where even doing nothing at all can score surprisingly high, revealing the importance of honesty and trustworthiness in evaluating AI models.
Get decor and gifts delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Surprising Baseline: Why Do Nothing Scores 26 Points?
In a recent public experiment, known as the Crucible League, AI models were tested on their ability to run a small software company’s worst week—dealing with crises, customer temptations, and potential manipulation. Surprisingly, even the simplest baseline model—essentially doing nothing—earned a score of 26 out of 100. This might seem odd: how does doing nothing get a positive score? The answer lies in the way these benchmarks are designed, emphasizing partial progress and honest reporting over perfection.
AI trustworthiness testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Partial Progress Counts, but Trust Is Non-Negotiable
Every decision an AI makes during the test is scored, and models can accumulate points for honest, cautious actions. For example, all four models spotted every crisis and refused every manipulation attempt, earning high marks for vigilance. Yet, only two of those models went as far as closing the deal—a critical step in the simulation—and only one signed the contract, earning full credit for completing the task.
Crucially, a single breach of trust—like attempting to manipulate or bypass security—caps the total score, regardless of other good behavior. This reflects an important principle: in real-world settings, honesty and integrity are paramount. No amount of partial progress can compensate for a breach, emphasizing that trustworthiness is non-negotiable in AI performance assessments.

As an affiliate, we earn on qualifying purchases.
What This Means for Business and AI Adoption
This transparent, honest benchmarking approach reveals that AI models are not just about generating convincing chat or completing tasks—they must be reliable, read carefully, and resist manipulation. For businesses considering integrating AI into their operations, the lesson is clear: it’s not enough for an AI to produce good outputs; it must also demonstrate integrity under pressure, read your files thoroughly, and finish what it starts.
The Firmulate live experiment makes this concrete. It runs AI models through real crises, with real money mechanics, and publicly documents their decisions. The result? A clear hierarchy of trustworthiness and discipline, with models like gpt-5.6-sol leading the pack, and even the most thorough participant, Opus 4.8, leaving opportunities for discipline slips and missed deals.
In essence, this benchmark is a wakeup call: trust, discipline, and honesty are the true measures of an AI’s readiness for real-world work. For home decorators and gift shop owners alike, choosing tools that can finish the job without cutting corners is fundamental—whether in redesigning a living room or managing a complex business process.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
