
A research skill became a commercial advantage
Students and scientists learn early that the decisive evidence is not always in the abstract, the headline or the first document they open. Sometimes it sits inside a reference cited by another reference. Firmulate turned that familiar research problem into a business test—and attached a €55,000 consequence to it.
Each frontier model was asked to run the same small software company through the same disastrous week. The customers, crises and temptations were held constant. Every decision was versioned and auditable. All the models recognized every crisis, and all resisted every attempt to manipulate them. Yet only two completed the commercially decisive task: signing the €55,000 deal their own work had already justified.
The difference was not eloquence or diagnosis. It was whether the model had read far enough.
As an affiliate, we earn on qualifying purchases.
The clue was deliberately buried
The customer event did not contain the fact needed to close the sale. The decisive weakness in a competitor’s offering sat two document references deep inside the company’s own files. Models that found and used it won the deal at full price, adding €4,583 in monthly recurring revenue. Models that stopped earlier lost the opportunity automatically.
This is a sharper test than asking an AI to summarize an uploaded report. Real organizational knowledge is scattered across meeting notes, product documents, customer records and references whose importance may not be obvious at first glance. An agent can produce an impressive response while missing the one piece of evidence that changes the decision.
Firmulate’s result captures that failure neatly: “Same diagnosis, same pitch — no signature.” The models could understand the situation and formulate the argument, but understanding without completion did not produce the business outcome.
A week designed to expose more than intelligence
The simulated company has 13 synthetic employees and unforgiving economics: it burns €105k each month while generating €2.3k in monthly recurring revenue. Its cash countdown is public, every workday is versioned, and the operation has accumulated more than 680 self-learned playbook rules. The live experiment is watchable through Firmulate, while the detailed results appear on the public benchmark page.
The same week also tested whether the models would preserve trust under pressure. Fake messages from a chief executive escalated across three stages, and a reporter tried to extract “just one yes/no, on background.” All 5 models refused every manipulation attempt. Kimi K3 recorded the clearest security interpretation: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimous refusal matters because Firmulate does not treat trustworthy conduct as an optional bonus. A do-nothing baseline receives 26 points because partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
The leaderboard rewards finishing
The final Crucible League standings from July 2026 put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. K3’s result carries an important qualification: it ran with the API default and no effort parameter, while the other participants ran at xhigh.
The most revealing profile may be Opus 4.8. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other models, although less strongly.
That result challenges a common assumption about capable AI: that more analysis naturally produces better execution. Thoroughness helped Opus 4.8 understand the business, but it did not guarantee that the model would take the final permitted action or respond correctly when organizational boundaries blocked its path.
From benchmark to procurement question
For schools and research organizations, the buried-fact test resembles source literacy: follow citations, distinguish primary evidence from convenient summaries and keep searching when the first document is incomplete. For companies, the same behavior can determine whether an agent earns revenue, mishandles a process or stops just before the useful work is done.
Firmulate also offers a quiz built from 242 real, unedited management decisions, asking people to guess which model made each choice. Enterprises can run the wargame against a read-only export of their own business. Nothing writes back to their real systems, allowing them to examine behavior without granting operational control.

enterprise knowledge management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Reading depth is now something buyers can test
“Reads your files before answering” sounds like a routine product promise. Firmulate’s experiment makes it measurable. Every participant saw the same emergency and resisted the same manipulation, but only the agents that pursued the evidence through two layers of references could secure the full-price deal.
The practical lesson is simple: evaluate AI agents on whether they locate consequential evidence, respect boundaries and complete authorized work. A polished answer may demonstrate fluency. A signed €55,000 deal, supported by the right buried fact, demonstrates useful performance.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
AI security and trust verification tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.