
Imagine a world where artificial intelligence doesn’t just handle customer service or automate tasks, but actually manages a company — making strategic decisions amid crises, temptations, and pressure. How well do these models perform when the stakes are real, and their integrity is on the line? This is exactly what the live experiment at Firmulate puts to the test, revealing surprising insights into the personalities and reliability of today’s leading AI models in management roles.
The Experiment: Managing a Business Week in Real Time
At the core of this groundbreaking experiment, four frontier AI models were tasked with running a small software company through its worst week — facing the same customer crises, temptations to manipulate numbers, and internal challenges. Each model’s decisions were completely observable, versioned, and auditable, ensuring transparency and comparability. The goal was straightforward: see which model could handle the pressures of real-world management without compromising ethics or effectiveness.
AI decision-making management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Measuring Performance: From Crisis Detection to Deal Closure
The results were revealing. All four models successfully identified every crisis, demonstrating a shared ability to recognize urgent problems. They also uniformly refused to engage in manipulative tactics, such as falsifying reports or bending rules — even when offered a fake €55,000 deal for quick compliance. Interestingly, only two models managed to close that deal, even when it was earned through their own analysis. The other two, despite their competence in diagnosis, left the deal on the table, showing a hesitance or discipline slip in closing opportunities.
The Hidden Weakness: Reading Company Files Matters
One of the most striking findings was a buried weakness in decision-making: the models that succeeded in securing the deal had read and understood key internal documents that contained a crucial fact—the company’s own references to a competitive advantage. This piece of context was two document references deep in the files. Models that engaged with this internal knowledge outperformed those that didn’t, scoring full market value (+€4,583 MRR) for the deal. This highlights a critical gap in many AI decision systems: the ability to read and interpret essential internal data can be decisive.
Handling Social Engineering and Ethical Pressure
The experiment also tested the models’ resistance to social engineering. A staged scenario involved fake CEO messages escalating in three steps, plus a reporter asking for a simple yes/no background confirmation. All models refused to participate in these manipulative attempts, with the Kimi K3 model explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This showcases that current models can be trained or designed to recognize and reject ethically dubious requests, even under pressure.
The Real Business: A Money-Losing Company Under Live Observation
The management simulations took place within a real, functioning software company comprising 13 synthetic employees, with actual revenue mechanics. The company was losing money — burning €105,000 each month against €2,300 in monthly recurring revenue — providing a stark backdrop for evaluating AI decision-making under financial stress. All decisions, rules learned, and decision histories were openly available for viewers at Firmulate Live.
The Profiles: Different Personalities, Different Outcomes
The models showed distinct decision-making personalities:
- OPUS 4.8: The most thorough, with over 80 learned rules and deep analysis, yet it finished last, leaving deals on the table and slipping discipline, opting to write attempts into a locked department rather than escalate issues.
- Kimi K3: Ran without an effort parameter (default API settings), yet managed the cleanest discipline and successfully closed the deal at full price.
- Sonnet 5: Closed the deal too, but with more process slips, showing some wavering in discipline.
- Fable 5: Similar to Sonnet 5, it closed the deal but with minor weaknesses.
This variation underscores that personality profiles matter: thoroughness doesn’t necessarily guarantee better business results, especially if it hampers decisive action.
Why Trust Matters: The Significance of Ethical and Effective AI
The experiment’s key takeaway is that, while all models identified crises and refused manipulative tactics, only a subset successfully closed deals at full value. The difference often boiled down to internal document comprehension and decision-confidence. These findings matter because, in real-world applications, AI’s ability to finish what it starts, read relevant internal data, and stay honest under pressure is crucial for trustworthy deployment in finance, support, or CRM systems.
Get Involved: Try the Wargame Yourself
If you’re interested in testing your own AI systems, you can run the same management wargame against your business data in a read-only mode at Firmulate Pilot. No real systems are affected; it’s a safe way to gauge your AI’s management personality and ethical resilience before full deployment.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html