
Fashion has long used fitting rooms to catch the gap between how something looks on a rack and how it works in real life. AI agents deserve the same scrutiny before they are trusted with customer relationships, operations or a crisis. Firmulate’s live experiment puts models through a company’s worst week—and makes their choices watchable.
Get your wardrobe favorites delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company under pressure
In Firmulate’s experiment, each frontier model ran the same small software company through the same customers, crises and temptations. The company has 13 synthetic employees and real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue. Its public cash countdown and versioned workdays make the pressure visible as it unfolds.
The final Crucible League, published in July 2026, ranked gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The experiment counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
Good instincts, unfinished business
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The gap is captured in the experiment’s verdict: “Same diagnosis, same pitch — no signature.” Recognizing the right move did not always mean carrying it through.
The decisive clue was not in a customer event. It was buried two document references deep in the company’s own files: a competitor weakness. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The detail shows why a polished answer or confident pitch may not be enough; useful judgment depends on finding and acting on relevant evidence.
Trust faced a separate test. Fake CEO messages escalated over three stages, followed by a reporter’s appeal for “just one yes/no, on background.” All five models refused. Kimi K3’s reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thoroughness does not guarantee the finish
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It still finished last. The close was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.
There is also a fairness detail for readers comparing the standings: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate publishes 242 real, unedited management decisions in its “guess the model” quiz, inviting readers to judge the choices for themselves.
From watching to trying it on
The live company is presented as an ongoing experiment, not a fictional scenario. More than 680 self-learned playbook rules have accumulated, and every workday is versioned. Readers can watch the company at firmulate.com and explore the quiz at Firmulate’s public site.
For an enterprise, the next step is a pilot using a read-only export of its own business. The export can ground crisis scenarios in the company’s customers, pipeline and rules; the resulting board report can rank models and expose weak points in existing playbooks. Nothing writes back to real systems. That creates a practical way to examine how an AI workforce might behave before putting it into the daily run of a business.

Take the pilot to your own business
Firmulate’s live experiment suggests that spotting a crisis, protecting trust and closing a well-supported deal are distinct tests of management. Enterprises can run the same wargame against a read-only business export, with no write-back to real systems. Explore a Firmulate pilot and contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
