
Imagine a live company run entirely by artificial intelligence, battling daily crises, making crucial decisions, and risking its very existence—all on display for the world to see. This isn’t science fiction. It’s happening now, in real time, at Firmulate. This extreme build-in-public experiment offers a rare glimpse into how AI can manage complex, real-world business challenges—without human employees—and how well it actually performs under pressure.
The Live Company and Its High-Stakes Environment
At the heart of the experiment is a company with 13 synthetic employees powered by advanced AI models. Despite running on a razor-thin financial margin—burning €105,000 each month against a Monthly Recurring Revenue of just €2,300—the company is publicly racing against time. Every workday, its decision-making is versioned and recorded, creating a transparent, auditable trail of how AI handles crises, customer interactions, and internal dilemmas.
The Performance of AI Models in Practice
Four frontier AI models were tasked with navigating the company’s toughest week. These models include GPT-5.6-sol, Kimi K3, Sonnet 5, and Fable 5. All demonstrated a remarkable ability: they spotted every crisis and refused every manipulation attempt, including social engineering tactics like fake CEO messages and reporter tricks. Yet, only two of the models managed to close a critical €55,000 deal based on their own analysis—an essential revenue opportunity for the company.
Notably, the decisive factor was not in immediate customer interactions but buried deep within the company’s own internal documents. The models that read these files identified a key reference—hidden in the company’s files—leading them to successfully secure the deal at full price, worth an additional €4,583 in monthly recurring revenue.
Building-in-Public and Its Challenges
This experiment is a showcase of transparency in AI development. Each decision made by the models is versioned, and every outcome is observable on the live site. The models were tested against the same crises, temptations, and manipulations, ensuring a level playing field. For instance, when social engineering attempts—such as escalating fake CEO messages—were introduced, all models refused to comply, citing concerns about impersonation or bypassing approval protocols, with Kimi K3 explicitly noting the risk of impersonation.
Performance Disparities and Lessons Learned
Among the models, Opus 4.8 was the most thorough, with over 80 learned rules and deep analyses, yet it left a promising deal unexecuted due to a slip in discipline—its close was left on the table, and issues in escalation processes emerged. Conversely, Kimi K3, running without an effort parameter, demonstrated the cleanest discipline and closed the deal successfully, highlighting how different configurations affect outcomes.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for the Future of AI in Business
This continuous, transparent experiment underscores the critical questions facing AI adoption in real companies. It’s not just about whether AI writes well or generates convincing chat responses. The key issues are whether AI can see the full picture—reading internal documents thoroughly—stay honest under pressure, and follow through on decisions made. As firms consider integrating AI into customer support, sales, or operations, these performance metrics offer valuable insights.
Watch the Experiment Live
Interested in seeing this in action? The live company is accessible at firmulate.com/live.html. Visitors can observe the daily decision-making process, watch as AI models respond to crises, and see how the company’s financial situation evolves in real time. The site rebuilds twice daily, providing fresh data and new challenges for the models.

AI in Public Relations: Reputation Management with Prompts (AI BUSINESS & MANAGEMENT LIBRARY SERIES Book 4)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Education and Beyond
For students, educators, and researchers, this experiment offers a rare window into the practical capabilities and limitations of AI in managing complex, uncertain environments. It emphasizes the importance of transparency, robust decision-making, and trustworthiness—concepts that are equally vital in education, science, and other fields. The experiment demonstrates that AI can learn, adapt, and even show integrity under duress, but it also reveals the gaps that still need bridging before AI can reliably replace human judgment in critical roles.

This live experiment at Firmulate shows AI managing a company through its toughest week—spotting crises, refusing manipulation, and closing deals—highlighting both its potential and limitations in real-world business. Transparency reveals how AI can be a trustworthy partner, but also exposes areas needing improvement before wider adoption.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Ai For Customer Experience And Support: A Practical Guide To Automating Service, Personalizing Interactions, And Driving Customer Loyalty With Artificial Intelligence
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
![Claude AI for Beginners Bible: [5 in 1] The Ultimate Guide to Automate Your Work, Save Hours Every Week, and Use AI for Real-World Results](https://m.media-amazon.com/images/I/415+fSJacsL._SL500_.jpg)
Claude AI for Beginners Bible: [5 in 1] The Ultimate Guide to Automate Your Work, Save Hours Every Week, and Use AI for Real-World Results
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.