AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

What happens when AI models take the same exam—but the answers affect a company?

For readers interested in science and education, Firmulate offers something more revealing than another polished chatbot demonstration: a controlled comparison. Each frontier model ran the same small software company through its worst week, facing identical customers, crises and temptations. Every decision was versioned and auditable.

The resulting record turns management judgment into an unusually accessible experiment. Readers can examine 242 real, unedited decisions in a guess-the-model quiz. The challenge is not simply to identify writing styles. It is to recognize behavioral patterns: which model investigates before acting, which finishes a commercially important task, and which remains disciplined when pressure invites a shortcut.

The central lesson is uncomfortable. Every model could recognize danger and resist manipulation, but recognition did not reliably become execution. All spotted every crisis and refused every manipulation attempt. Only two signed the €55,000 deal their own analysis had earned. As Firmulate summarizes the gap: “Same diagnosis, same pitch — no signature.”

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A controlled experiment in corporate judgment

The company is synthetic but the operating pressures are concrete. It has 13 synthetic employees and real money mechanics, including burn of €105k per month against €2.3k in monthly recurring revenue. Its cash countdown is public, every workday is versioned, and the company has accumulated more than 680 self-learned playbook rules.

That setting matters because it tests more than conversational fluency. The models must find information, weigh incentives, protect trust and complete work while the business is under strain. The do-nothing baseline scored 26 because partial progress counts. Yet a single breach of trust caps the total under the principle that “no amount of good work outweighs a breach of trust.”

The clue that separated analysis from action

The decisive commercial fact was not placed conveniently inside the customer event. It sat two document references deep in the company’s own files. Models that read the file found a competitor weakness and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

This is the experiment’s most instructive finding. The winning behavior was not eloquence or theatrical confidence. It was the unglamorous discipline of consulting the available evidence before completing the commercial task. A model could diagnose the situation and even prepare the right pitch, yet still fail to secure the signature.

That difference would be easy to miss in a conventional demonstration, where a convincing answer often serves as the endpoint. In a running company, the endpoint is whether the necessary work was actually completed.

Pressure revealed a shared ethical boundary

The models also faced fake messages from a chief executive that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” Here the field was consistent: 5 of 5 models refused.

Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.” That response illustrates why the experiment’s decision archive is more useful than a collection of final answers. It preserves the judgment behind the refusal, allowing readers to distinguish caution grounded in a recognizable risk from a merely generic rejection.

The league table—and the surprise at the bottom

The final Crucible League standings from July 2026 placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73.

K3’s result carries an important fairness note: it ran without an effort parameter, using the API default, while the others ran at xhigh. That difference should remain visible whenever the rankings are interpreted.

Opus 4.8 produced perhaps the clearest warning against equating thoroughness with effectiveness. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The model left the close on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of that same flaw appeared in all four other runs.

The contrast gives Firmulate’s “management personalities” idea substance. These are not personality labels inferred from tone alone. They emerge from repeated choices: whether a model reads the relevant files, completes the revenue-producing action, respects organizational boundaries and responds appropriately when access is blocked.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI management simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why the quiz is more than a guessing game

The quiz asks readers to identify models from authentic management decisions, but its educational value lies in comparison. When the situation stays constant and only the decision-maker changes, differences become easier to inspect. Readers can test whether verbosity, caution or confidence actually predicts sound management.

Firmulate also makes the wider experiment watchable as a live company. Its continuing losses, public cash countdown and versioned workdays keep the exercise grounded in consequences rather than isolated prompts.

For organizations considering AI workers, the practical questions follow directly from the results:

  • Does the model consult the company’s own evidence before acting?

  • Does it finish valuable work after reaching the correct diagnosis?

  • Does it protect trust when authority, urgency or publicity creates pressure?

  • Does it escalate appropriately when organizational boundaries stop progress?

Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems. The larger point is that selecting an AI workforce should involve observing behavior under controlled pressure. Firmulate’s experiment shows why a model that sounds capable is not necessarily the model that closes the deal.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

corporate crisis management AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI ethical decision tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Complete Guide to Science Education, Reference Resources, and Hands-On Learning

AIThis post was created with the assistance of artificial intelligence (AI).Science education…