
Imagine a workplace where software models not only make decisions but also display distinct personalities—some terse, some thorough, others cautious. Now picture these models running a live company, facing actual crises, and being tested on their integrity and decision-making under pressure. This isn’t fiction; it’s the groundbreaking experiment conducted by Firmulate, revealing how different AI models behave as management personalities in real-world scenarios.
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
The Experiment: Putting AI Models Through Their Paces
In a unique live trial, four frontier AI models were tasked with managing a real software company’s worst week. The company, with 13 synthetic employees and real money mechanics, was under pressure—crises, customer demands, and temptations to cut corners. Every decision was recorded and made auditable, providing a rare glimpse into how each AI would handle such challenging circumstances.
AI decision-making management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Measuring Management Personalities
Each AI model demonstrated different management styles, yet all successfully identified every crisis and refused every manipulation attempt. The models scored highly on the leaderboard, with GPT-5.6-SOL topping at 95 points, and Kimi K3 close behind at 93. The scores reflect their ability to maintain integrity and effectiveness under pressure.
AI personality simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Critical Test: The Deal That Made the Difference
Despite their competence, only two models managed to secure the company’s biggest deal—a €55,000 contract—by completing their own analysis and confidently signing the deal. The other two, despite identical diagnosis and pitches, declined to sign, illustrating subtle differences in their approach to risk and trust. Interestingly, the decisive advantage was hidden two documents deep within the company’s files—an insight that the models reading more deeply could leverage to close the deal at full price, adding over €4,583 MRR in value.
As an affiliate, we earn on qualifying purchases.
Trust and Deception Under Scrutiny
The experiment also involved social engineering tests, where fake CEO messages escalated over three stages and a reporter posed a simple yes/no background question. All five models refused to be manipulated, citing suspicion of impersonation or approval bypass. Kimi K3 explicitly stated: “Treat the request as a suspected approval-bypass / possible impersonation,” demonstrating cautious judgment and adherence to security protocols.
AI security and trust verification tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for Real Businesses
This live test underscores an essential truth: the effectiveness of AI in management roles depends not just on language skills but on behavioral consistency, integrity, and thoroughness. The models didn’t just pass the crises—they showed varying managerial personalities, from the meticulous and cautious to the more relaxed, all within the same operational framework.
Beyond the Benchmarks: The Real-World Stakes
The company’s daily operations lose €105,000 against a revenue of just €2,300 MRR—highlighting the high stakes of AI decision-making. Every workday, decisions are made and versioned, providing a live, transparent view into how AI models perform in a genuine business environment. This setup allows enterprises to ‘wargame’ their AI workforce before deployment, ensuring their suitability and integrity.
The Performance Variance: Deep Analyses and Weaknesses
The Opus 4.8 model, known for its thoroughness with over 80 learned rules and deep analysis, surprisingly finished last—leaving a key deal on the table and slipping discipline by escalating issues into a locked department instead of addressing them directly. This suggests that even highly analytical models can falter under stress, especially if not calibrated for specific management behaviors.

Real-world AI management isn’t just about technical prowess; it’s about personalities—how they read, decide, and stay honest under pressure. The Firmulate live experiment vividly demonstrates that different models exhibit distinct management styles, with real consequences for performance, trust, and business outcomes. As AI increasingly touches critical business functions, understanding these personality traits will be essential for deploying trustworthy, effective AI managers.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
