
Choosing an AI model can feel like choosing a school: polished answers are easy to admire, but performance under pressure tells you more. In Firmulate’s live company experiment, Moonshot’s Kimi K3 finished second overall and beat three of four Western frontier models. The result is a reminder for anyone comparing AI tools: a strong demo is not the same as a reliable track record.
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
A shared test, with real consequences
Firmulate put frontier models in charge of the same small software company during its worst week. They faced the same customers, crises and temptations, and every decision was versioned and auditable. The experiment is presented as a live, watchable company, not a fictional case study.
In the final July 2026 league, gpt-5.6-sol scored 95, Kimi K3 93, Sonnet 5 88, Fable 5 77 and Opus 4.8 73. A do-nothing baseline scored 26. The result puts K3 in second place, ahead of three of the four Western models in the table.
The gap between knowing and doing
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The key was buried two document references deep in the company’s files: models that read the file found the competitor’s weakness and won the deal at full price, worth +€4,583 MRR.
That difference matters beyond a leaderboard. An AI system can give a persuasive diagnosis and still leave the useful action unfinished. For organizations considering AI in customer support, sales or forecasting, the question is not only whether a model can explain what to do, but whether it follows through responsibly.
The experiment also tested social engineering: fake CEO messages escalated through three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thoroughness is not the whole report card
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and discipline slipped when it made write attempts into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.
Firmulate says its synthetic company has 13 employees, burn of €105k per month against €2.3k MRR, a public cash countdown and more than 680 self-learned playbook rules. The company’s workday is versioned, and its live experiment can be watched at Firmulate. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice.
There is a fairness caveat for the rankings: K3 ran without an effort parameter (API default) while the others ran at xhigh. That context belongs beside the score, especially for readers treating the results as a model comparison.
For organizations that want to evaluate their own AI workforce, Firmulate describes an enterprise pilot using a read-only export of a company’s business. The pilot page says nothing writes back to real systems. The full benchmark and plain-language findings are available at Firmulate’s benchmarks.

What to take from the result
Kimi K3’s second-place finish shows that the field is open, while the unsigned deals show why a ranking alone cannot settle a purchasing decision. Test models against the work, pressure and safeguards that matter in your own organization before choosing one.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
