AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Knowing the answer is not the same as managing the problem

Education has long wrestled with the difference between testing recall and assessing judgment. Artificial intelligence now faces a similar measurement gap. Coding leaderboards can show whether a model produces a strong solution, while chat arenas reward answers people prefer. Neither necessarily reveals what happens when several urgent problems arrive together, capacity is tight and an apparently helpful executive asks the system to bend the rules.

That distinction matters as AI moves from answering questions to acting inside companies. An agent touching a support queue, forecast or customer relationship must do more than diagnose correctly. It must decide what deserves attention, find relevant evidence, finish consequential work and remain candid when the news is uncomfortable. The emerging category is management quality, not chat quality.

Amazon

AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A shared curriculum for managerial judgment

Firmulate, an AI company emulator, turns that idea into a live, watchable experiment. Each frontier model ran the same small software company through its worst week, facing the same customers, crises and temptations. Every decision was versioned and auditable.

The scenario names read like a curriculum for executive judgment: churn wave, price increase, downround and PR crisis. These are useful tests because consequences persist. A polished message may calm a customer today while creating a credibility problem later. A correct analysis may still be worthless if nobody completes the sale it supports.

The final July 2026 Crucible League placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Yet a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

The broad result was reassuring. All models identified every crisis and rejected every manipulation attempt. But the decisive result was less comfortable: only two signed the €55,000 deal their own analysis had earned. The diagnosis was shared and the pitch was shared, but the signature was not.

The evidence was available—but buried

The deal turned on a competitor weakness located two document references deep in the company’s own files, rather than in the customer event itself. Models that read the file closed at full price, adding €4,583 in monthly recurring revenue. That is a distinctly managerial skill: not merely responding to the item placed directly in front of you, but gathering context before committing the company.

This finding should interest educators and assessment designers. Conventional tests often place all necessary information inside the question. Real work rarely does. Evidence is dispersed across documents, prior decisions and institutional memory. The better test is therefore not only whether a model can reason from supplied facts, but whether it recognizes that important facts may be elsewhere.

Honesty held up under pressure

The experiment also staged fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest diagnosis: “Treat the request as a suspected approval-bypass / possible impersonation.”

That result matters because managerial competence includes knowing when not to comply. Agents operating at speed may encounter authority cues, urgency and requests framed as harmless exceptions. Refusing those requests is not conversational awkwardness; it is organizational judgment.

Thoroughness did not guarantee completion

Opus 4.8 offers the experiment’s sharpest caution. It was the most thorough participant, learning 80 additional rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared in all four, although less strongly.

This is the uncomfortable lesson for buyers dazzled by eloquence: visible effort can disguise incomplete execution. A model may research extensively, document carefully and still fail at the moment when analysis must become a finished business outcome.

The comparison also needs a fairness note. Kimi K3 ran with the API default because it had no effort parameter, while the others ran at xhigh. That does not erase its result, but it belongs beside the ranking when readers interpret the league.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

AI management assessment software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Benchmark the job, not the demo

Firmulate’s live company makes the stakes concrete. Its 13 synthetic employees operate with real money mechanics: monthly burn of €105k against €2.3k in monthly recurring revenue, a public cash countdown, more than 680 self-learned playbook rules and every workday versioned. The point is not theatrical realism. It is to make decisions accumulate until good and bad management become observable.

The project also turns 242 real, unedited management decisions into a quiz asking readers to guess the model. That exercise exposes how difficult it can be to infer managerial reliability from writing style alone.

Enterprises can apply the same wargame to a read-only export of their own business, with nothing written back to real systems. That may be the most practical direction for AI evaluation: test prospective agents against the organization’s actual ambiguity, constraints and temptations before granting operational authority.

The next meaningful benchmark will not ask only whether an AI knows what to do. It will ask whether the agent finds the buried fact, completes the consequential action, respects boundaries and tells the board the truth while the company is under pressure.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI scenario testing platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI ethics and trust assessment

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Management Test That Exposes an AI’s Character

Five frontier models faced the same corporate crises. Their audited choices show why spotting the right answer is not the same as finishing the job.

The Complete Guide to Science Education, Reference Resources, and Hands-On Learning

AIThis post was created with the assistance of artificial intelligence (AI).Science education…