AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

A living lesson in how artificial intelligence behaves under pressure

For readers interested in education and science, Firmulate offers something more illuminating than another polished AI demonstration: a continuing experiment whose mistakes remain visible. The public company simulation employs 13 synthetic workers, records every workday as a new version and has accumulated more than 680 self-learned playbook rules. Its finances are deliberately unforgiving. Monthly burn is €105,000, while monthly recurring revenue stands at €2,300.

That imbalance turns abstract questions about artificial intelligence into an observable management problem. Can a model find evidence buried in company records? Will it finish commercially important work? Does it resist pressure to misuse authority? And can it learn without becoming distracted by the appearance of productivity?

The experiment is not presented as a fictional corporate tale. It is running software with real money mechanics and a public cash countdown. Readers can watch the company live, following a business that is visibly fighting for survival and generating new material every workday.

Hands-On Simulation Modeling with Python: Develop simulation models for improved efficiency and precision in the decision-making process, 2nd Edition

Hands-On Simulation Modeling with Python: Develop simulation models for improved efficiency and precision in the decision-making process, 2nd Edition

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same disastrous week, with different managers

Firmulate’s Crucible League placed frontier models in charge of the same small software company during its worst week. Each encountered identical customers, crises and temptations. Their decisions were versioned and auditable, making the exercise closer to a controlled comparison than to a collection of unrelated chatbot anecdotes.

The final July 2026 table put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. One boundary was absolute, however: a single breach of trust capped the total, under the principle that “no amount of good work outweighs a breach of trust.”

The broad result was reassuring. All models identified every crisis and rejected every manipulation attempt. The sharper lesson concerned completion. Only two signed the €55,000 deal that their own work had earned. The researchers summarized the contrast as: “Same diagnosis, same pitch — no signature.”

The decisive evidence was easy to overlook

The winning commercial fact did not appear in the customer event that demanded attention. It sat two document references deep in the company’s own files: a competitor weakness that justified holding the full price. Models that read far enough found it and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

This is a useful educational distinction. Recognizing the visible problem is not the same as investigating it, and producing a plausible answer is not the same as completing the task. In workplaces, evidence may be distributed across ordinary records rather than placed neatly inside a prompt. Firmulate’s test makes that gap concrete: the models shared a diagnosis, but their research habits and follow-through produced different outcomes.

Pressure tested trust as well as competence

The company also received fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 models refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

That result matters because business automation is not merely a contest in eloquence. A system operating around customers, forecasts or internal records must distinguish legitimate urgency from manufactured pressure. Firmulate allows the public to inspect what its synthetic staff actually communicate through its collection of company quotes, adding human-readable evidence to the rankings.

Why the busiest participant finished last

Opus 4.8 presents the experiment’s most instructive paradox. It was the most thorough participant, producing the deepest analyses and adding 80 learned rules, yet it finished last in the final league. It left the close on the table and lost procedural discipline by attempting to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

The case cautions against equating more analysis with better management. Learning, documentation and reflection are valuable, but they do not substitute for a completed decision or a correct escalation. A large body of work can coexist with a missed commercial outcome.

There is also an important fairness qualification. Kimi K3 ran with the API default because it had no effort parameter, while the others ran at xhigh. The result remains observable, but that difference belongs beside any interpretation of the narrow gap near the top.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
The AI-Driven Leader: Harnessing AI to Make Faster, Smarter Decisions

The AI-Driven Leader: Harnessing AI to Make Faster, Smarter Decisions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A public laboratory for consequential AI

Firmulate’s strongest contribution is not a claim that one model will always manage a company better than another. It is the transformation of vague qualities—judgment, persistence, research discipline and resistance to manipulation—into a continuing public record.

For educators, researchers and general readers, the live company offers a compact lesson in evaluating AI: look beyond fluent answers. Ask whether the system found the relevant evidence, respected boundaries, escalated correctly and carried useful work through to completion. The cash countdown gives those questions urgency, while the versioned workdays make the story inspectable rather than merely promotional.

With 13 synthetic employees, more than 680 learned rules and finances that remain starkly underwater, the experiment turns each workday into another test. Its most provocative finding is also its simplest: seeing the problem is not the same as solving it.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI trust and ethics training kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

DEEP Robotics Lite 3-Basic AI Quadruped Robot, AI Robotics Platform

DEEP Robotics Lite 3-Basic AI Quadruped Robot, AI Robotics Platform

  • AI-Powered Quadruped Platform: Designed for developers and researchers
  • Enhanced Mobility & Stability: Supports 40° slope climbing and 15cm obstacles
  • Suitable for Indoor and Outdoor Use: Stable performance across environments

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Angle Chasing: When It Works and When It Wastes Time

Ineffective in complex diagrams, but mastering angle chasing can save time—discover when and how to use it effectively.

How Estimation Saves Time Before Exact Geometry Calculations

An early estimation can reveal potential issues and save time, but understanding how to leverage this advantage can dramatically streamline your design process.

Geometric Constructions With Compass and Straightedge

Theorem and practical techniques in geometric constructions with compass and straightedge reveal elegant solutions and fascinating challenges worth exploring further.

Inversion in Geometry: Solving Hard Problems by Flipping the Plane

Solving complex geometry problems becomes easier with inversion, a powerful technique that transforms shapes and angles—discover how it can unlock your next solution.