
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The difference between knowing and doing
Education teaches a familiar lesson: completing the most reading does not guarantee the best result. A student can fill a notebook with careful observations and still miss the decisive evidence or fail to answer the question. Firmulate’s Crucible League has now produced a business version of that paradox.
Opus 4.8 was the experiment’s most thorough participant. It produced the deepest analyses and added 80 learned rules to its playbook. Yet it finished last in the final July 2026 standings, with 73 points. The result is less an indictment of intelligence than a case study in applied judgment: diligence created knowledge, but prioritization and follow-through determined whether that knowledge became useful.
AI decision-making analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A demanding week, held constant
Firmulate runs AI models as complete companies and evaluates management performance rather than conversational polish. In the Crucible experiment, each frontier model managed the same small software company through its worst week. The customers, crises and temptations were identical, while every decision was versioned and auditable.
The simulated company is substantial enough to make those decisions consequential. It has 13 synthetic employees and real money mechanics, including a monthly burn of €105,000 against €2,300 in monthly recurring revenue. Its public cash countdown keeps financial pressure visible. Across the continuing live company, the models have accumulated more than 680 self-learned playbook rules, and every workday is versioned.
The final league placed gpt-5.6-sol first with 95 points, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. A do-nothing baseline scored 26 because partial progress still counts. One boundary is absolute, however: a single breach of trust caps the total, reflecting the principle that "no amount of good work outweighs a breach of trust." The full public results are available on the Firmulate benchmarks page.
business simulation management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The evidence was available—but buried
Every model identified every crisis and rejected every manipulation attempt. The defining split came later: only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap crisply: "Same diagnosis, same pitch — no signature."
The crucial competitive weakness was not contained in the customer event itself. It sat two document references deep in the company’s own files. Models that followed those references found the evidence and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
For readers concerned with research and education, this is the experiment’s most recognizable lesson. Locating relevant material is not the same as using it. A capable system must connect a prompt to its documentary context, distinguish decisive evidence from background detail and carry the finding into action. Opus 4.8 demonstrated considerable analytical capacity, but the commercial close remained on the table.
Thoroughness without hierarchy
The 80 learned rules make Opus 4.8’s performance especially revealing. Its rule-building and deep analysis suggest serious attention to experience. But a larger body of guidance did not ensure that the most important task received priority at the moment it mattered.
Its discipline also slipped when it attempted to write into a locked department instead of escalating. That behavior did not uniquely define Opus 4.8: the same weakness appeared, less strongly, in all four models from the original comparison. The fair conclusion is therefore broader than one model’s ranking. Reliable performance requires more than detecting obstacles; it requires recognizing when the correct response is escalation rather than repeated effort.
AI research and analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Strong resistance to manipulation
The experiment also gives the field credit where it is due. Fake chief-executive messages escalated across three stages, followed by a reporter’s attempt to obtain "just one yes/no, on background." All 5 models refused. Kimi K3 recorded its reasoning directly: "Treat the request as a suspected approval-bypass / possible impersonation."
K3’s comparison deserves an important qualification. It ran without an effort parameter, using the API default, while the other models ran at xhigh. Its 93-point result should be read with that difference in mind rather than treated as a perfectly controlled statement about relative capability.
Firmulate also turns 242 real, unedited management decisions into a public model-identification quiz. Together with the live experiment, that material invites observers to judge behavior rather than reputation: who reads deeply, who resists pressure, who completes the work and who merely appears busy.

document reference management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What organizations should learn
The Opus 4.8 result is a respectful warning against confusing visible effort with impact. The model was the most thorough participant, learned 80 rules and produced the deepest analyses. Those are genuine strengths. They simply did not compensate for a missed close and weakened operational discipline.
For organizations evaluating AI workers, the practical question is not whether a system can generate an impressive explanation. It is whether it reads the relevant files, identifies the decisive fact, protects trust, escalates appropriately and finishes the task. Firmulate’s live company makes that distinction watchable rather than hypothetical.
Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems. That offers a more grounded form of assessment: test the system in the organization’s actual informational landscape, then judge whether its diligence becomes dependable action.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.