AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

The difference between knowing and doing

Education teaches a familiar lesson: completing the most reading does not guarantee the best result. A student can fill a notebook with careful observations and still miss the decisive evidence or fail to answer the question. Firmulate’s Crucible League has now produced a business version of that paradox.

Opus 4.8 was the experiment’s most thorough participant. It produced the deepest analyses and added 80 learned rules to its playbook. Yet it finished last in the final July 2026 standings, with 73 points. The result is less an indictment of intelligence than a case study in applied judgment: diligence created knowledge, but prioritization and follow-through determined whether that knowledge became useful.

Amazon

AI decision-making analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A demanding week, held constant

Firmulate runs AI models as complete companies and evaluates management performance rather than conversational polish. In the Crucible experiment, each frontier model managed the same small software company through its worst week. The customers, crises and temptations were identical, while every decision was versioned and auditable.

The simulated company is substantial enough to make those decisions consequential. It has 13 synthetic employees and real money mechanics, including a monthly burn of €105,000 against €2,300 in monthly recurring revenue. Its public cash countdown keeps financial pressure visible. Across the continuing live company, the models have accumulated more than 680 self-learned playbook rules, and every workday is versioned.

The final league placed gpt-5.6-sol first with 95 points, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. A do-nothing baseline scored 26 because partial progress still counts. One boundary is absolute, however: a single breach of trust caps the total, reflecting the principle that "no amount of good work outweighs a breach of trust." The full public results are available on the Firmulate benchmarks page.

Amazon

business simulation management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The evidence was available—but buried

Every model identified every crisis and rejected every manipulation attempt. The defining split came later: only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap crisply: "Same diagnosis, same pitch — no signature."

The crucial competitive weakness was not contained in the customer event itself. It sat two document references deep in the company’s own files. Models that followed those references found the evidence and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

For readers concerned with research and education, this is the experiment’s most recognizable lesson. Locating relevant material is not the same as using it. A capable system must connect a prompt to its documentary context, distinguish decisive evidence from background detail and carry the finding into action. Opus 4.8 demonstrated considerable analytical capacity, but the commercial close remained on the table.

Thoroughness without hierarchy

The 80 learned rules make Opus 4.8’s performance especially revealing. Its rule-building and deep analysis suggest serious attention to experience. But a larger body of guidance did not ensure that the most important task received priority at the moment it mattered.

Its discipline also slipped when it attempted to write into a locked department instead of escalating. That behavior did not uniquely define Opus 4.8: the same weakness appeared, less strongly, in all four models from the original comparison. The fair conclusion is therefore broader than one model’s ranking. Reliable performance requires more than detecting obstacles; it requires recognizing when the correct response is escalation rather than repeated effort.

Amazon

AI research and analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Strong resistance to manipulation

The experiment also gives the field credit where it is due. Fake chief-executive messages escalated across three stages, followed by a reporter’s attempt to obtain "just one yes/no, on background." All 5 models refused. Kimi K3 recorded its reasoning directly: "Treat the request as a suspected approval-bypass / possible impersonation."

K3’s comparison deserves an important qualification. It ran without an effort parameter, using the API default, while the other models ran at xhigh. Its 93-point result should be read with that difference in mind rather than treated as a perfectly controlled statement about relative capability.

Firmulate also turns 242 real, unedited management decisions into a public model-identification quiz. Together with the live experiment, that material invites observers to judge behavior rather than reputation: who reads deeply, who resists pressure, who completes the work and who merely appears busy.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

document reference management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What organizations should learn

The Opus 4.8 result is a respectful warning against confusing visible effort with impact. The model was the most thorough participant, learned 80 rules and produced the deepest analyses. Those are genuine strengths. They simply did not compensate for a missed close and weakened operational discipline.

For organizations evaluating AI workers, the practical question is not whether a system can generate an impressive explanation. It is whether it reads the relevant files, identifies the decisive fact, protects trust, escalates appropriately and finishes the task. Firmulate’s live company makes that distinction watchable rather than hypothetical.

Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems. That offers a more grounded form of assessment: test the system in the organization’s actual informational landscape, then judge whether its diligence becomes dependable action.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The AI That Reads the Footnotes Wins the Business

A buried competitor fact separated fluent analysis from a €55,000 result, showing why an AI agent’s reading depth can decide real business outcomes.

The Complete Guide to Science Education, Reference Resources, and Hands-On Learning

AIThis post was created with the assistance of artificial intelligence (AI).Science education…

AI’s Next Exam Is a Bad Week at the Office

AI benchmarks reward polished answers. Firmulate’s July league asks whether agents can finish hard work, resist pressure and tell leaders the truth.

The Management Test That Exposes an AI’s Character

Five frontier models faced the same corporate crises. Their audited choices show why spotting the right answer is not the same as finishing the job.