firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

In the rapidly evolving world of artificial intelligence, how do we measure whether these digital workers are truly reliable? A recent public experiment by Firmulate sheds light on what it really means for an AI to be trustworthy — and why sometimes, doing nothing is the best score of all.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

Understanding the Benchmark: More Than Just Scores

At first glance, AI performance might seem straightforward: the higher the score, the better the model. But in a groundbreaking live experiment, Firmulate put four frontier models through a simulated week of managing a small software company facing crises, customer manipulations, and ethical dilemmas. The results challenge simplistic notions and highlight the importance of trust, thoroughness, and discipline in AI decision-making.

The Experiment in a Nutshell

  • All four models managed the same company scenario, with identical crises and temptations.
  • They were tasked to identify issues, respond ethically, and secure a deal valued at €55,000.
  • Decisions were fully versioned and auditable, ensuring transparency.

Key Findings: Do Nothing Is Not Always Bad

Surprisingly, all four models detected every crisis and refused every manipulation attempt, demonstrating robust ethical behavior. Yet, only two managed to close the deal at full value, while the others fell short, leaving some opportunities on the table. This resulted in a score of 26 for a ‘do-nothing’ baseline, which might seem low but actually reflects a fundamental truth about AI behavior.

Amazon

AI trustworthiness testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Lesson: Reading Beyond the Surface

The decisive advantage for the top performers was not in superficial decision-making but in reading deeper into the company’s documentation. The winning models identified critical information buried two document references deep—information that was essential to closing the deal at full price. Those who read more thoroughly—and thus acted more accurately—won the business.

Ethical and Manipulation Challenges

The models faced social engineering attacks, including fake CEO messages staged over multiple steps and a reporter trick. Impressively, all five models refused to be manipulated, citing concerns about impersonation or approval bypasses. This demonstrates a vital feature for AI in real-world applications: the capacity to recognize and resist deception.

Why Trust and Discipline Matter More Than Scores

In the experiment, a strict rule was applied: a single breach of trust caps the overall score. For example, if an AI attempts to escalate issues improperly or bypass protocols, it cannot recover its reputation in the scoring. This ‘floor’ at 26 points for doing nothing underscores that in high-stakes environments, partial progress isn’t enough—trustworthiness is non-negotiable.

What This Means for Business AI

Many people focus on chat quality or speed when evaluating AI, but these metrics miss the bigger picture. The real question is: does the AI finish what it starts, read the relevant information thoroughly, and stay honest under pressure? These are the qualities that determine whether AI can be a reliable partner in managing business operations, especially in contexts involving sensitive data and ethical considerations.

The Firmulate Live Platform: Transparent and Watchable

Firmulate’s live experiment is accessible online, allowing anyone to see the models in action within a simulated company. The platform manages real money mechanics—burning €105k per month against a modest €2.3k monthly recurring revenue, with over 680 self-learned rules guiding decision-making. Every decision is versioned and auditable, providing a transparent window into how AI agents behave under pressure.

Deep Dive into Model Performance

The Opus 4.8 model, though the most thorough in rule-learning, finished last—leaving the deal on the table and slipping into unprofessional escalation procedures. Similarly, Kimi K3, which ran at a default API setting, scored just one point shy of the top, illustrating how configuration choices impact discipline and performance.

Implications for the Future of AI in Business

This experiment underscores a core truth: AI systems must be evaluated not only on their ability to generate convincing outputs but on their capacity for ethical, thorough, and disciplined decision-making. A ‘do-nothing’ baseline score of 26 shows that minimal activity isn’t the goal; rather, it’s the quality, integrity, and trustworthiness of actions that count.

Takeaway

For businesses considering AI adoption, the key takeaway is clear: look beyond superficial metrics. Prioritize models that can read deeply, remain disciplined under pressure, and refuse unethical shortcuts. Trust isn’t built by clever chatter but by consistent, honest performance—something a simple baseline score can reveal.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Best Native Trees for Wildlife Uk

AIThis post was created with the assistance of artificial intelligence (AI).It is…

Which Is a Feature of Australian Aboriginal Depictions of the Natural World?

AIThis post was created with the assistance of artificial intelligence (AI).Are you…

Indigenous Room Ideas

AIThis post was created with the assistance of artificial intelligence (AI).When it…

Biggest React Native Apps

AIThis post was created with the assistance of artificial intelligence (AI).When discussing…