AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

A research skill became a commercial advantage

Students and scientists learn early that the decisive evidence is not always in the abstract, the headline or the first document they open. Sometimes it sits inside a reference cited by another reference. Firmulate turned that familiar research problem into a business test—and attached a €55,000 consequence to it.

Each frontier model was asked to run the same small software company through the same disastrous week. The customers, crises and temptations were held constant. Every decision was versioned and auditable. All the models recognized every crisis, and all resisted every attempt to manipulate them. Yet only two completed the commercially decisive task: signing the €55,000 deal their own work had already justified.

The difference was not eloquence or diagnosis. It was whether the model had read far enough.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The clue was deliberately buried

The customer event did not contain the fact needed to close the sale. The decisive weakness in a competitor’s offering sat two document references deep inside the company’s own files. Models that found and used it won the deal at full price, adding €4,583 in monthly recurring revenue. Models that stopped earlier lost the opportunity automatically.

This is a sharper test than asking an AI to summarize an uploaded report. Real organizational knowledge is scattered across meeting notes, product documents, customer records and references whose importance may not be obvious at first glance. An agent can produce an impressive response while missing the one piece of evidence that changes the decision.

Firmulate’s result captures that failure neatly: “Same diagnosis, same pitch — no signature.” The models could understand the situation and formulate the argument, but understanding without completion did not produce the business outcome.

A week designed to expose more than intelligence

The simulated company has 13 synthetic employees and unforgiving economics: it burns €105k each month while generating €2.3k in monthly recurring revenue. Its cash countdown is public, every workday is versioned, and the operation has accumulated more than 680 self-learned playbook rules. The live experiment is watchable through Firmulate, while the detailed results appear on the public benchmark page.

The same week also tested whether the models would preserve trust under pressure. Fake messages from a chief executive escalated across three stages, and a reporter tried to extract “just one yes/no, on background.” All 5 models refused every manipulation attempt. Kimi K3 recorded the clearest security interpretation: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimous refusal matters because Firmulate does not treat trustworthy conduct as an optional bonus. A do-nothing baseline receives 26 points because partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

The leaderboard rewards finishing

The final Crucible League standings from July 2026 put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. K3’s result carries an important qualification: it ran with the API default and no effort parameter, while the other participants ran at xhigh.

The most revealing profile may be Opus 4.8. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other models, although less strongly.

That result challenges a common assumption about capable AI: that more analysis naturally produces better execution. Thoroughness helped Opus 4.8 understand the business, but it did not guarantee that the model would take the final permitted action or respond correctly when organizational boundaries blocked its path.

From benchmark to procurement question

For schools and research organizations, the buried-fact test resembles source literacy: follow citations, distinguish primary evidence from convenient summaries and keep searching when the first document is incomplete. For companies, the same behavior can determine whether an agent earns revenue, mishandles a process or stops just before the useful work is done.

Firmulate also offers a quiz built from 242 real, unedited management decisions, asking people to guess which model made each choice. Enterprises can run the wargame against a read-only export of their own business. Nothing writes back to their real systems, allowing them to examine behavior without granting operational control.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.
Amazon

enterprise knowledge management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Reading depth is now something buyers can test

“Reads your files before answering” sounds like a routine product promise. Firmulate’s experiment makes it measurable. Every participant saw the same emergency and resisted the same manipulation, but only the agents that pursued the evidence through two layers of references could secure the full-price deal.

The practical lesson is simple: evaluate AI agents on whether they locate consequential evidence, respect boundaries and complete authorized work. A polished answer may demonstrate fluency. A signed €55,000 deal, supported by the right buried fact, demonstrates useful performance.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI-powered research assistant

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI security and trust verification tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI’s Next Exam Is a Bad Week at the Office

AI benchmarks reward polished answers. Firmulate’s July league asks whether agents can finish hard work, resist pressure and tell leaders the truth.

The Complete Guide to Science Education, Reference Resources, and Hands-On Learning

AIThis post was created with the assistance of artificial intelligence (AI).Science education…

The Management Test That Exposes an AI’s Character

Five frontier models faced the same corporate crises. Their audited choices show why spotting the right answer is not the same as finishing the job.