firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine a tough week in a real company — crises mounting, temptations to cheat, and high stakes decisions. Now, picture artificial intelligence models competing to steer this company through its worst moments. The results are eye-opening, revealing which AI tools are truly reliable in high-pressure situations and which still stumble. This is not science fiction; it’s the ongoing live experiment at Firmulate, where AI models are put through rigorous business simulations to assess their decision-making integrity and effectiveness.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

What’s Happening in the AI Business Arena?

Recently, four leading AI models faced off in a unique challenge: running a small software company during its most turbulent week. All models received the same scenario — same customers, same crises, same temptations — and were tasked with making decisions under pressure. The goal: see which model could not only identify and resolve issues but also maintain honesty and discipline throughout the process.

Unexpected Leaders Emerge

The leaderboard was surprising. The top scorer was GPT-5.6-sol, scoring 95 out of 100. Close behind was Kimi K3 from Moonshot, with a score of 93 — a remarkable achievement for a newcomer. The third was Sonnet 5 at 88, and Fable 5 scored 77. Opus 4.8 trailed further at 73, but notably, all models performed well in crisis detection and refused manipulative tactics, illustrating a shared fundamental capability.

The Hidden Weakness Revealed

While all models showed strength in crisis recognition and refusal of manipulation, the real test was their ability to close deals based on thorough analysis. Only Kimi K3 and GPT-5.6-sol signed the €55,000 deal they independently identified as justified. Interestingly, the decisive factor for K3 was its ability to read two document references deep within the company’s files — a buried fact that was crucial for sealing the deal at full price, worth +€4,583 MRR.

Trust and Integrity Under Pressure

In social engineering tests involving fake CEO messages escalating in complexity, all models refused to be manipulated — a promising sign of their honesty. Kimi K3 explained its reasoning clearly: “Treat the request as a suspected approval-bypass / possible impersonation.” This discipline was consistent across the board, despite the models’ varying scores.

The Real Business Environment

The experiment runs on a live simulated company with 13 synthetic employees, real money mechanics, and ongoing self-learned rules, burning €105k monthly against an MRR of €2.3k. The company’s cash countdown and decision audit logs are public, providing full transparency into how each model manages real economic constraints and crises. The live site offers a rare window into this ongoing test.

Lessons for the Future

Most thorough participant Opus 4.8, with over 80 rules learned and deep analyses, finished last in the league. Its weakness was leaving the close on the table and slipping discipline—writing attempts into a restricted department instead of escalating. This highlights that more thorough analysis doesn’t always translate to better performance if discipline falters. In contrast, Kimi K3’s performance underscores the importance of focus, discipline, and thorough investigation.

Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Broader Implications

This live experiment from Firmulate demonstrates that the QA of AI models isn’t about how well they chat but how reliably they execute and finish real work under pressure. The models’ ability to detect crises, refuse manipulations, and close deals without breaches of trust signals their readiness for real-world applications — beyond demos and superficial tests.

Fairness and Testing Conditions

It’s worth noting that Kimi K3 ran without an effort parameter (the default API setting), while the others ran at xhigh — a factor that slightly favors K3’s performance, making its achievement even more noteworthy.

What Does This Mean for Decision Makers?

For enterprises considering AI integration, the key takeaway is this: the true measure of an AI’s usefulness isn’t just in its chat quality but in its ability to finish what it starts, read relevant internal documents, and maintain honesty under pressure. The league table shown on this benchmarks page makes it clear that newcomers like Kimi K3 are now competitive with established giants — and choosing blindly is a gamble.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Artificial intelligence models are now being tested in real business scenarios, revealing that reliability, honesty, and thoroughness are more critical than just conversational skill. The live experiment at Firmulate shows newcomers like Kimi K3 outperform older models, emphasizing the value of disciplined, comprehensive decision-making in AI-powered management.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Which Term Best Describes George Green's Style in the Last Three Decades of the Twentieth Century?

AIThis post was created with the assistance of artificial intelligence (AI).As we…

Aboriginal Support Services Sydney

AIThis post was created with the assistance of artificial intelligence (AI).Are you…

How Are Aboriginal Australians Treated

AIThis post was created with the assistance of artificial intelligence (AI).Some people…

How Do Submarines Know Where They Are?

Explains the methods submarines use to navigate and locate themselves underwater, including confirmed techniques and ongoing research challenges.