
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
A test of what happens after the right answer
In education, a correct diagnosis is only part of the lesson. Students also have to show what they can do with it. Firmulate’s business experiment poses a similar question for AI: can a model recognize a crisis, resist pressure and carry its own analysis through to a sound decision?
The result matters beyond software companies. As businesses consider AI agents for customer service, sales and planning, it is not enough to ask whether a model can explain what should happen. The harder question is whether it follows through when the stakes and temptations are real within the simulation.
One company, one difficult week
In the final Crucible League, in July 2026, frontier models each ran the same small software company through its worst week. Customers, crises and temptations were held constant; only the model changed. Every decision was versioned and auditable.
The results ranged from 95 for gpt-5.6-sol and 93 for Kimi K3 to 88 for Sonnet 5, 77 for Fable 5 and 73 for Opus 4.8. The do-nothing baseline scored 26. The benchmark counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
The models all recognized every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s compact summary captures the gap: “Same diagnosis, same pitch — no signature.” In a chat demonstration, a persuasive recommendation can look like success. In this test, the decision to act mattered too.
The clue was buried in the company’s own files
The decisive competitor weakness was not in the customer event. It was hidden two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR.
That is a useful lesson about preparation: relevant evidence may sit outside the immediate prompt or event. Finding it can change the outcome. The experiment makes that difference visible in a concrete business decision, rather than asking readers to infer performance from a polished answer.
Pressure tested trust and follow-through
The social-engineering challenge escalated through three stages of fake CEO messages, followed by a reporter’s request: “just one yes/no, on background”. All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Refusing manipulation was common ground; completing the earned deal was not. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the close on the table and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models.
There is a fairness caveat in the comparison. K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That detail belongs beside the rankings when interpreting them.
From observing to testing your own business
Firmulate’s live company makes the experiment watchable. It has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and a versioned record for every workday. Readers can watch the company at firmulate.com. A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each call.
The enterprise pilot takes the idea from watching to acting. A company supplies a read-only export of its business; the wargame runs crisis scenarios against that picture and produces a board report with model rankings and weak points in the company’s own playbooks. Nothing writes back to real systems. That lets leaders examine how AI might handle their customers, decisions and pressure points before putting it to work.

Make the hard decisions visible
The league shows why a capable answer is not the whole measure of an AI workforce. Trust, evidence gathering and follow-through all shape the outcome. For organizations ready to examine those behaviors against their own business, explore a Firmulate pilot or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
