firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine hiring a new home decorator who claims to transform your space. But before they even start, they show up with a messy toolkit, uncertain about your style, and occasionally even break a vase. Would you trust them to deliver? In AI testing, a similar story unfolds—where even doing nothing at all can score surprisingly high, revealing the importance of honesty and trustworthiness in evaluating AI models.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get decor and gifts delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Surprising Baseline: Why Do Nothing Scores 26 Points?

In a recent public experiment, known as the Crucible League, AI models were tested on their ability to run a small software company’s worst week—dealing with crises, customer temptations, and potential manipulation. Surprisingly, even the simplest baseline model—essentially doing nothing—earned a score of 26 out of 100. This might seem odd: how does doing nothing get a positive score? The answer lies in the way these benchmarks are designed, emphasizing partial progress and honest reporting over perfection.

Amazon

AI trustworthiness testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Partial Progress Counts, but Trust Is Non-Negotiable

Every decision an AI makes during the test is scored, and models can accumulate points for honest, cautious actions. For example, all four models spotted every crisis and refused every manipulation attempt, earning high marks for vigilance. Yet, only two of those models went as far as closing the deal—a critical step in the simulation—and only one signed the contract, earning full credit for completing the task.

Crucially, a single breach of trust—like attempting to manipulate or bypass security—caps the total score, regardless of other good behavior. This reflects an important principle: in real-world settings, honesty and integrity are paramount. No amount of partial progress can compensate for a breach, emphasizing that trustworthiness is non-negotiable in AI performance assessments.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.
Amazon

AI benchmarking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What This Means for Business and AI Adoption

This transparent, honest benchmarking approach reveals that AI models are not just about generating convincing chat or completing tasks—they must be reliable, read carefully, and resist manipulation. For businesses considering integrating AI into their operations, the lesson is clear: it’s not enough for an AI to produce good outputs; it must also demonstrate integrity under pressure, read your files thoroughly, and finish what it starts.

The Firmulate live experiment makes this concrete. It runs AI models through real crises, with real money mechanics, and publicly documents their decisions. The result? A clear hierarchy of trustworthiness and discipline, with models like gpt-5.6-sol leading the pack, and even the most thorough participant, Opus 4.8, leaving opportunities for discipline slips and missed deals.

In essence, this benchmark is a wakeup call: trust, discipline, and honesty are the true measures of an AI’s readiness for real-world work. For home decorators and gift shop owners alike, choosing tools that can finish the job without cutting corners is fundamental—whether in redesigning a living room or managing a complex business process.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI integrity assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI model evaluation kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Health Sciences Surges In Global Coverage

Health sciences are experiencing a surge in worldwide media coverage, with 22 mentions in recent monitoring, signaling increased interest and focus.

It’s Getting Harder To Focus Every Day

Recent studies indicate a growing struggle with focus among adults, raising concerns about mental health and productivity in daily life.

The Science Behind Why We Love To Drive

Exploring the psychological and neurological reasons behind why many people find driving enjoyable, based on recent trend signals and scientific insights.

AI’s Integrity Holds Firm Under Threat: A Real-World Test of Trust and Security

A real-world AI experiment tested models against social engineering threats; all refused manipulation, showing integrity can be secured before deployment, not after the breach.