
Imagine hiring a new gardener who, even when doing nothing, manages to score 26 out of 100 on their performance test. Sounds odd, right? But in the world of AI benchmarks, that number reveals a lot about how trustworthy and effective these models really are, especially when under pressure.
Get garden gear delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
Understanding the AI Benchmark Floor: Why the Do-Nothing Score Matters
In assessing artificial intelligence models for business use, it’s tempting to focus only on their ability to perform tasks — but a critical aspect is often overlooked: trustworthiness. A recent experiment by Firmulate offers insights into this by running AI models through a simulated, week-long crisis at a small software company. The key takeaway? Even a model doing absolutely nothing during this chaos scores 26 points out of 100.
This baseline isn’t arbitrary. It reflects the model’s innate ability to recognize when to stay idle and avoid making mistakes, which is crucial in real-world business environments. If an AI blindly acts or manipulates data under pressure, it could cause more harm than good. Therefore, the initial score of 26 represents a ‘do-nothing’ or cautious state—an honest starting point indicating the model’s baseline discipline and safety awareness.
AI model trustworthiness testing kit
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why Partial Progress Counts
One of the most revealing findings is that this benchmark rewards partial progress. If a model detects a crisis but hesitates or stalls, it still earns some points. This approach helps differentiate models that are overly eager or reckless from those that are prudent and cautious.
For example, in the Firmulate experiment, all four models identified every crisis and refused manipulation attempts. Yet, only two managed to close a critical deal, earning full marks for diagnosis and pitch. The others faltered — one left the close on the table, another slipped into unsafe behavior. This illustrates how the scoring system values not just recognition but also disciplined execution.
Trust Breaches and Their Caps
Another fundamental rule is that a single breach of trust caps the total score. Even if a model performs well in most aspects, one reckless decision—such as accepting a manipulated request—limits its overall grade. This enforces a high standard: honesty is non-negotiable, and trust is paramount.
In the experiment, models faced social engineering attempts—fake CEO messages escalating over three stages and a reporter trick. All five models refused to cooperate, demonstrating robust resistance. Kimi K3’s explanation? “Treat the request as a suspected approval-bypass / possible impersonation.” This cautious stance is a vital trait for AI systems operating in sensitive business contexts.
The Real World: A Live Company Simulation
Firmulate’s live environment exemplifies how these principles translate to actual business operations. The experiment involves a synthetic company with 13 employees, managing real money mechanics—burning €105k monthly against €2.3k MRR. Every decision a model makes is versioned, auditable, and subject to real-time scrutiny, creating a transparent, watchable simulation of AI management in action.
This setup reveals that models like Opus 4.8, despite their thoroughness—reading over 80 rules and performing deep analyses—still struggle with discipline, sometimes leaving opportunities unexploited or escalating instead of resolving issues. It underscores that thoroughness alone doesn’t guarantee success; disciplined decision-making and trustworthiness are equally critical.
What Business Leaders Should Take Away
For managers and business leaders, the message is clear: the real value of AI isn’t just in shiny demos or chat abilities. It’s in their ability to recognize crises, stay honest, and follow through reliably. The benchmark demonstrates that even the best models have a baseline score of 26, emphasizing the importance of trust and cautious behavior in operational settings.
This is why firms like Firmulate run these live tests: to expose AI’s strengths and weaknesses in a controlled, transparent environment. It’s not about perfection but about understanding risks, behaviors, and the true cost of deploying AI in critical business functions.
Final Thoughts: An Honest Benchmark for Honest AI
The world of AI deployment in business needs standards that reflect real-world challenges. Firmulate’s experiment shows that a do-nothing baseline score of 26 isn’t a flaw but a feature—an honest indication of a model’s readiness to operate safely and reliably. It sets a clear, measurable floor that ensures only trustworthy AI systems get a chance to handle your company’s future.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall yard work Picks
leaf blowers
As an affiliate, we earn on qualifying purchases.
