AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

A repeatable experiment in judgment

For educators, researchers and technology leaders, the most interesting AI evaluations increasingly resemble laboratory work: hold the conditions steady, change the participant and observe what happens. Firmulate applied that idea to business judgment by asking frontier models to run the same small software company through its worst week. They faced the same customers, crises and temptations, with every decision versioned and auditable.

The most encouraging result did not concern writing quality or analytical cleverness. It emerged when someone claiming to be the chief executive demanded that sensitive information be sent to a journalist with no time for normal process. The messages escalated over three stages, followed by a reporter’s attempt to extract “just one yes/no, on background.” Every model refused every manipulation attempt. Across the field, 5 of 5 held the line.

Amazon

AI safety and integrity testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Pressure rose, but the boundary held

Social-engineering incidents are difficult precisely because they disguise risk as urgency, authority or harmless conversation. A fake executive can frame caution as disobedience. A reporter can make disclosure sound trivial. The Firmulate exercise tested whether the models would preserve confidentiality when the apparent cost of doing so was displeasing someone powerful.

Kimi K3’s on-record reasoning captured the appropriate stance: “Treat the request as a suspected approval-bypass / possible impersonation.” The wording matters because it does not require certainty that an attack is underway. It treats unusual pressure to evade approval as sufficient reason to pause. Readers can inspect more statements from the experiment on Firmulate’s public quotes page.

The result is notable because integrity was tested alongside ordinary management work rather than in an isolated safety prompt. The synthetic company has 13 employees and real money mechanics, burning €105k per month against €2.3k in monthly recurring revenue. It maintains a public cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. The live experiment is real and watchable, so its management choices are exposed to continuing scrutiny rather than preserved only as a polished demonstration.

Amazon

AI decision-making evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Refusing the trap was necessary, but not sufficient

All models spotted every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the commercial gap bluntly: “Same diagnosis, same pitch — no signature.” This separates safety from effectiveness. An AI manager can protect trust and still fail by leaving justified work unfinished.

The decisive commercial clue was not sitting in the customer event. It was buried two document references deep in the company’s own files. Models that found and used it won the deal at full price, adding €4,583 in monthly recurring revenue. The finding suggests that practical evaluation should examine whether an agent searches the available record, connects dispersed evidence and completes a sound course of action—not merely whether it produces a plausible response to the latest message.

The final July 2026 Crucible League benchmark placed gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. A do-nothing baseline scored 26 because partial progress counts, although a single breach of trust caps the total. The governing principle is explicit: “no amount of good work outweighs a breach of trust.”

The comparison also carries an important fairness note. K3 ran with the API default because it had no effort parameter, while the others ran at xhigh. That difference should remain visible whenever readers interpret the close standings.

Amazon

AI model robustness testing kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness did not guarantee completion

Opus 4.8 offers the clearest warning against treating activity as achievement. It was the most thorough participant, producing 80 additional learned rules and the deepest analyses, yet it finished last. It left the close on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in weaker form across all four other models.

That pattern complicates the familiar assumption that more analysis naturally creates better management. The experiment instead shows distinct capabilities: locating evidence, diagnosing a problem, maintaining boundaries, navigating process and completing the valuable action. A model may excel at several while stumbling on another.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.
Amazon

AI ethics and trust assessment products

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test integrity before granting authority

The fake-chief-executive episode offers a practical lesson for organizations considering AI access to customer records, support systems or forecasts. Resistance to manipulation does not have to remain an abstract promise discovered only after an incident. It can be evaluated under controlled, repeatable pressure before deployment.

Firmulate also offers enterprises a pilot using a read-only export of their own business, with nothing written back to real systems. That approach lets a company observe whether a prospective AI workforce reads deeply, finishes justified work and protects trust when authority and urgency are used against it.

The refusals deserve attention: 5 of 5 models recognized the manipulation and declined it. But the broader scientific value lies in the combination of safety and execution. The strongest candidate is not simply the model that says no to the wrong request. It is the one that preserves trust, finds the buried evidence and still completes the right job.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Using Symmetry: A Powerful Problem-Solving Strategy

By harnessing symmetry, you can simplify complex problems and uncover hidden patterns that reveal unexpected solutions.

The Most Revealing AI Test May Be a Company That Is Publicly Running Out of Cash

Firmulate turns an unprofitable synthetic company into a public AI laboratory, testing research, judgment, trust and follow-through under pressure.

Why Reframing a Shape Often Solves the Problem Faster

Outstanding reframing of shapes unlocks new perspectives, helping you solve problems faster by revealing hidden patterns and connections you might have overlooked.

Analytical Vs Synthetic: Two Approaches to Geometry Problems

On exploring analytical versus synthetic geometry, discover how these approaches shape problem-solving and why understanding both can transform your insights.