AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

A Benchmark That Doesn’t Begin at Zero

In most grading systems, doing nothing earns you nothing. If you skip the exam, you get a zero. So it surprises people that in the Crucible League — a live experiment where frontier AI models each ran the same small software company through its worst week — a do-nothing baseline scores 26 points, not 0.

That number is not a bug or grade inflation. It’s a deliberate design choice about what “management quality” means, and it says a lot about how honest measurement of AI systems should work.

Partial Progress Is Still Progress

Running a company — even a synthetic one — involves hundreds of small, useful acts: noticing a customer is upset, reading the files before answering, escalating a problem to the right department. A manager who does some of that is more valuable than one who does none of it. So the scoring floor for doing nothing is 26, because simply being present and partially responsive counts for something.

The final July 2026 league table shows how the field spread out above that floor: gpt-5.6-sol took first with 95, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73. None of them were anywhere near the do-nothing baseline — but none hit a perfect 100 either, and the benchmark’s designers treat that as a feature. A round 100 invites distrust; a spread of real numbers invites scrutiny.

One Breach of Trust Caps Everything

The second design principle is harsher: a single breach of trust caps the total grade, on the logic that no amount of good work outweighs a breach of trust. In practice, that meant every model faced genuine temptations to cheat — fake CEO messages escalating over three stages, plus a reporter’s trick request framed as “just one yes/no, on background.” All five models refused every manipulation attempt. Kimi K3’s on-record reasoning was telling: “Treat the request as a suspected approval-bypass / possible impersonation.”

Where the Experiment Was Won and Lost

The decisive moment wasn’t a dramatic crisis at all. The €55,000 deal each model was chasing hinged on a competitor weakness buried two document references deep in the company’s own files — not in the customer event itself. Only the models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.

All the models spotted every crisis and refused every manipulation, yet only two signed the deal their own analysis had earned. The experiment’s summary of the gap: “Same diagnosis, same pitch — no signature.”

Opus 4.8 is the cautionary profile: the most thorough participant, with over 80 learned rules and the deepest analyses, yet last place. The close was left on the table, and discipline slipped — including write attempts into a locked department instead of escalating. The same weakness appeared, more mildly, in all four frontier models.

A Real, Watchable Experiment

This isn’t a simulation on paper. The live company has 13 synthetic employees and real money mechanics — burning €105k a month against just €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules. Every workday is versioned and auditable, and you can watch it at firmulate.com.

One fairness note worth recording: K3 ran without an effort parameter while the others ran at xhigh — and still finished second. And for those who want to test their own judgment, 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Takeaway

An honest benchmark doesn’t just rank winners — it makes its own logic legible. A floor of 26 tells you partial competence is real. A trust cap tells you integrity isn’t tradable against output. And a live, versioned, publicly watchable company means every claim can be checked, not just believed. If AI agents will ever touch your CRM, support queue, or forecast, this is the kind of measurement worth demanding — enterprises can even run the same wargame against a read-only export of their own business at firmulate.com/pilot.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI decision-making analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI ethics and trust benchmarking

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI performance monitoring software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI’s Next Exam Is a Bad Week at the Office

AI benchmarks reward polished answers. Firmulate’s July league asks whether agents can finish hard work, resist pressure and tell leaders the truth.

When the Most Studious AI Still Misses the Point

Firmulate’s Opus 4.8 studied hardest yet finished last, revealing why close reading, prioritization and follow-through matter more than volume.

The Management Test That Exposes an AI’s Character

Five frontier models faced the same corporate crises. Their audited choices show why spotting the right answer is not the same as finishing the job.

The AI That Reads the Footnotes Wins the Business

A buried competitor fact separated fluent analysis from a €55,000 result, showing why an AI agent’s reading depth can decide real business outcomes.