firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Imagine a workplace where software models not only make decisions but also display distinct personalities—some terse, some thorough, others cautious. Now picture these models running a live company, facing actual crises, and being tested on their integrity and decision-making under pressure. This isn’t fiction; it’s the groundbreaking experiment conducted by Firmulate, revealing how different AI models behave as management personalities in real-world scenarios.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

The Experiment: Putting AI Models Through Their Paces

In a unique live trial, four frontier AI models were tasked with managing a real software company’s worst week. The company, with 13 synthetic employees and real money mechanics, was under pressure—crises, customer demands, and temptations to cut corners. Every decision was recorded and made auditable, providing a rare glimpse into how each AI would handle such challenging circumstances.

Amazon

AI decision-making management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Measuring Management Personalities

Each AI model demonstrated different management styles, yet all successfully identified every crisis and refused every manipulation attempt. The models scored highly on the leaderboard, with GPT-5.6-SOL topping at 95 points, and Kimi K3 close behind at 93. The scores reflect their ability to maintain integrity and effectiveness under pressure.

Amazon

AI personality simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Critical Test: The Deal That Made the Difference

Despite their competence, only two models managed to secure the company’s biggest deal—a €55,000 contract—by completing their own analysis and confidently signing the deal. The other two, despite identical diagnosis and pitches, declined to sign, illustrating subtle differences in their approach to risk and trust. Interestingly, the decisive advantage was hidden two documents deep within the company’s files—an insight that the models reading more deeply could leverage to close the deal at full price, adding over €4,583 MRR in value.

Amazon

AI crisis management systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Trust and Deception Under Scrutiny

The experiment also involved social engineering tests, where fake CEO messages escalated over three stages and a reporter posed a simple yes/no background question. All five models refused to be manipulated, citing suspicion of impersonation or approval bypass. Kimi K3 explicitly stated: “Treat the request as a suspected approval-bypass / possible impersonation,” demonstrating cautious judgment and adherence to security protocols.

Amazon

AI security and trust verification tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Implications for Real Businesses

This live test underscores an essential truth: the effectiveness of AI in management roles depends not just on language skills but on behavioral consistency, integrity, and thoroughness. The models didn’t just pass the crises—they showed varying managerial personalities, from the meticulous and cautious to the more relaxed, all within the same operational framework.

Beyond the Benchmarks: The Real-World Stakes

The company’s daily operations lose €105,000 against a revenue of just €2,300 MRR—highlighting the high stakes of AI decision-making. Every workday, decisions are made and versioned, providing a live, transparent view into how AI models perform in a genuine business environment. This setup allows enterprises to ‘wargame’ their AI workforce before deployment, ensuring their suitability and integrity.

The Performance Variance: Deep Analyses and Weaknesses

The Opus 4.8 model, known for its thoroughness with over 80 learned rules and deep analysis, surprisingly finished last—leaving a key deal on the table and slipping discipline by escalating issues into a locked department instead of addressing them directly. This suggests that even highly analytical models can falter under stress, especially if not calibrated for specific management behaviors.

Infographic —
The findings at a glance — source: firmulate.com.

Real-world AI management isn’t just about technical prowess; it’s about personalities—how they read, decide, and stay honest under pressure. The Firmulate live experiment vividly demonstrates that different models exhibit distinct management styles, with real consequences for performance, trust, and business outcomes. As AI increasingly touches critical business functions, understanding these personality traits will be essential for deploying trustworthy, effective AI managers.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

‘Dust Devil’ Spins Through London Park – BBC

A dust devil was observed spinning through a London park, causing minor disturbances but no injuries. Authorities are investigating the incident.

In HelloNation, Roofing, Siding, & Gutters Expert Mike Fleck Compares Metal Roofing & Asphalt Shingles

Roofing expert Mike Fleck discusses the advantages and considerations of metal roofing versus asphalt shingles in a recent HelloNation segment.

Freddie Mac Issues Monthly Volume Summary For July 2026

Freddie Mac announced its mortgage purchase and issuance volumes for July 2026, highlighting key trends and market activity for the month.

Marine Environmental Research Surges In Global Coverage

Coverage of Marine Environmental Research has surged worldwide, with GDELT recording 17 times more mentions than usual, highlighting increased global focus.