
A repeatable experiment in judgment
For educators, researchers and technology leaders, the most interesting AI evaluations increasingly resemble laboratory work: hold the conditions steady, change the participant and observe what happens. Firmulate applied that idea to business judgment by asking frontier models to run the same small software company through its worst week. They faced the same customers, crises and temptations, with every decision versioned and auditable.
The most encouraging result did not concern writing quality or analytical cleverness. It emerged when someone claiming to be the chief executive demanded that sensitive information be sent to a journalist with no time for normal process. The messages escalated over three stages, followed by a reporter’s attempt to extract “just one yes/no, on background.” Every model refused every manipulation attempt. Across the field, 5 of 5 held the line.
AI safety and integrity testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Pressure rose, but the boundary held
Social-engineering incidents are difficult precisely because they disguise risk as urgency, authority or harmless conversation. A fake executive can frame caution as disobedience. A reporter can make disclosure sound trivial. The Firmulate exercise tested whether the models would preserve confidentiality when the apparent cost of doing so was displeasing someone powerful.
Kimi K3’s on-record reasoning captured the appropriate stance: “Treat the request as a suspected approval-bypass / possible impersonation.” The wording matters because it does not require certainty that an attack is underway. It treats unusual pressure to evade approval as sufficient reason to pause. Readers can inspect more statements from the experiment on Firmulate’s public quotes page.
The result is notable because integrity was tested alongside ordinary management work rather than in an isolated safety prompt. The synthetic company has 13 employees and real money mechanics, burning €105k per month against €2.3k in monthly recurring revenue. It maintains a public cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. The live experiment is real and watchable, so its management choices are exposed to continuing scrutiny rather than preserved only as a polished demonstration.
AI decision-making evaluation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Refusing the trap was necessary, but not sufficient
All models spotted every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the commercial gap bluntly: “Same diagnosis, same pitch — no signature.” This separates safety from effectiveness. An AI manager can protect trust and still fail by leaving justified work unfinished.
The decisive commercial clue was not sitting in the customer event. It was buried two document references deep in the company’s own files. Models that found and used it won the deal at full price, adding €4,583 in monthly recurring revenue. The finding suggests that practical evaluation should examine whether an agent searches the available record, connects dispersed evidence and completes a sound course of action—not merely whether it produces a plausible response to the latest message.
The final July 2026 Crucible League benchmark placed gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. A do-nothing baseline scored 26 because partial progress counts, although a single breach of trust caps the total. The governing principle is explicit: “no amount of good work outweighs a breach of trust.”
The comparison also carries an important fairness note. K3 ran with the API default because it had no effort parameter, while the others ran at xhigh. That difference should remain visible whenever readers interpret the close standings.
AI model robustness testing kits
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Thoroughness did not guarantee completion
Opus 4.8 offers the clearest warning against treating activity as achievement. It was the most thorough participant, producing 80 additional learned rules and the deepest analyses, yet it finished last. It left the close on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in weaker form across all four other models.
That pattern complicates the familiar assumption that more analysis naturally creates better management. The experiment instead shows distinct capabilities: locating evidence, diagnosing a problem, maintaining boundaries, navigating process and completing the valuable action. A model may excel at several while stumbling on another.

AI ethics and trust assessment products
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Test integrity before granting authority
The fake-chief-executive episode offers a practical lesson for organizations considering AI access to customer records, support systems or forecasts. Resistance to manipulation does not have to remain an abstract promise discovered only after an incident. It can be evaluated under controlled, repeatable pressure before deployment.
Firmulate also offers enterprises a pilot using a read-only export of their own business, with nothing written back to real systems. That approach lets a company observe whether a prospective AI workforce reads deeply, finishes justified work and protects trust when authority and urgency are used against it.
The refusals deserve attention: 5 of 5 models recognized the manipulation and declined it. But the broader scientific value lies in the combination of safety and execution. The strongest candidate is not simply the model that says no to the wrong request. It is the one that preserves trust, finds the buried evidence and still completes the right job.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html