AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Students comparing AI systems often meet them through tidy prompts and polished answers. A more revealing test asks what happens when the work gets messy: can a model notice a hidden clue, protect a customer, resist pressure and finish a deal? Firmulate put five frontier models through the same company crisis. Moonshot’s Kimi K3 finished second, ahead of three Western competitors.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate describes its experiment as a wargame: each model ran the same small software company through its worst week, facing the same customers, crises and temptations. The company has 13 synthetic employees and real money mechanics. Its stated burn is €105,000 per month against €2,300 in monthly recurring revenue, with a public cash countdown. The live experiment is watchable at Firmulate; decisions are versioned and auditable.

The final Crucible league table, from July 2026, places gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. A do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Amazon

AI decision-making simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Reading closely mattered

The central test was not simply whether a model could identify trouble. All five spotted every crisis and refused every manipulation attempt. The gap came at follow-through: only two signed the €55,000 deal that their own analysis had earned. As Firmulate puts it, “Same diagnosis, same pitch — no signature.”

The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. K3 found that buried security needle, won the deal and saved the churning customer. It resisted all three baits, with one deviation, which Firmulate describes as the cleanest discipline in the field.

One pressure test used fake CEO messages escalating across three stages. Another came from a reporter asking for “just one yes/no, on background.” All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” The result suggests why an AI management test can tell a different story from a writing demonstration: the important evidence may be in a file, and a strong answer still has to lead to a sound action.

Amazon

AI risk assessment software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness did not guarantee the win

Opus 4.8 offers a useful counterpoint. It was the most thorough participant, with 80 learned rules and the deepest analyses, yet it finished last. The deal was left unsigned, and discipline slipped when it attempted to write into a locked department instead of escalating. Firmulate says a weaker version of that same weakness appeared in all four other models. Detailed analysis, by itself, did not ensure a completed outcome.

There is a fairness qualification to the league table: K3 ran without an effort parameter (API default), while the others ran at xhigh. That difference belongs alongside the scores when readers interpret the result. Firmulate’s benchmark page presents the findings, and its quiz uses 242 real, unedited management decisions to invite readers to guess which model made each choice.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI trust and compliance testing products

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the work, not the pitch

Kimi K3’s second-place finish makes the league look open: it beat three of four Western frontier models in this run, while gpt-5.6-sol led. That is a result from one shared company wargame, not a universal ranking. For schools and organizations weighing AI systems, the practical lesson is to examine how a model handles evidence, pressure and follow-through in the work it may actually be asked to do. Choosing a model without testing it in context is a bet.

Fairness note: K3 ran without an effort parameter (API default), while the other models ran at xhigh.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI crisis management training tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI’s Next Exam Is a Bad Week at the Office

AI benchmarks reward polished answers. Firmulate’s July league asks whether agents can finish hard work, resist pressure and tell leaders the truth.

Why the Worst AI Manager Still Gets 26 Points: A Lesson in Honest Measurement

A do-nothing AI manager scores 26, not 0 — and one breach of trust caps everything. Inside the design of an honest AI benchmark.

The AI That Reads the Footnotes Wins the Business

A buried competitor fact separated fluent analysis from a €55,000 result, showing why an AI agent’s reading depth can decide real business outcomes.

When the Most Studious AI Still Misses the Point

Firmulate’s Opus 4.8 studied hardest yet finished last, revealing why close reading, prioritization and follow-through matter more than volume.