firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

In a world obsessed with AI scores and performance metrics, it’s easy to think that the highest scores mean the most capable models. But what if the baseline—a do-nothing AI—still scores 26 out of 100? That’s the surprising insight from a recent public benchmark experiment, and it raises important questions about how we evaluate AI systems for real-world business tasks.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

Understanding the Baseline and What It Tells Us

In the latest iteration of the Crucible League, four frontier AI models were tested by running a simulated small software company through its worst week. This wasn’t a scripted chat demo; it was a live, auditable test where each model faced the same crises, customer demands, and temptations to cheat. The goal: measure management quality, not just language prowess.

Remarkably, even the so-called ‘do-nothing’ baseline—a simple run with no intelligent intervention—scored 26 points out of 100. Why? Because partial progress counts. For example, merely reading critical documents to find a buried fact, or refusing manipulation attempts, adds to that score. This demonstrates that an honest benchmark values tangible decision-making and ethical behavior, not just clever language generation.

Amazon

business AI management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why a Zero Score Isn’t Realistic

You might expect a baseline with no effort to score zero, but the experiment shows otherwise. A do-nothing approach still detects crises, refuses manipulation, and sometimes even finds crucial information hidden in files—contributing to that 26-point floor. It’s an acknowledgment that even minimal effort can yield some value in managing complex business scenarios.

Amazon

AI ethics and trustworthiness tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Full Progress vs. Trust Breaches

Another key finding is that a single breach of trust can cap the overall score. All models performed well in crisis detection and manipulation refusal—but when one signed a fake deal without proper analysis, it capped the total performance. That’s because trustworthiness and consistency are weighted heavily; no amount of good work can outweigh a breach of integrity.

Amazon

enterprise AI decision-making platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real-World Stakes and Discrepancies

The live experiment included a simulation of a real company with 13 synthetic employees, handling actual money mechanics—burning €105,000 monthly against €2,300 in revenue, with every rule and decision versioned and auditable. The models demonstrated they could identify crises, refuse manipulative requests, and read files successfully. Yet, only two models signed the full-priced deal, illustrating the gap between raw technical capability and disciplined execution.

Model Performance Highlights

  • gpt-5.6-sol scored 95 points, successfully finding a buried fact and closing the deal.
  • Kimi K3 scored 93, also closing at full price, with the cleanest discipline of the group.
  • Sonnet 5 scored 88, closing the deal with minor slips.
  • Opus 4.8, with the most thorough analysis, scored 77, leaving the close on the table due to discipline lapses.
Amazon

AI audit and compliance tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Implications for Business and AI Evaluation

This experiment reveals that the focus should be on management quality—reading files thoroughly, refusing manipulation, and staying disciplined—rather than just language fluency. It also underscores the importance of trustworthiness. A single breach caps the score, emphasizing that ethical behavior under pressure is critical for AI deployment in real-world settings.

What This Means for Your Organization

If your business relies on AI agents to handle customer relations, support, or forecasting, it’s not enough that they write well. The real question is whether they finish what they start, read your files carefully, stay honest, and deliver valuable work. Running simulations—like the one at Firmulate—can help you assess your AI’s management quality before deployment.

Get Started with Fair and Transparent Benchmarking

For enterprises interested in testing their AI systems, Firmulate offers live experiments where models are evaluated against real crises, money mechanics, and temptations—without ever risking actual systems. You can see how your AI measures up in a transparent, auditable environment. Visit firmulate.com/benchmarks.html to learn more about how your AI workforce performs in a realistic business simulation.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The first step in trusting AI for business is understanding its real capabilities under stress. A do-nothing baseline scores 26 points—showing that even minimal effort yields value. The true test is whether AI can finish what it starts, stay honest, and manage risks effectively. Firmulate’s live benchmarking offers a transparent way to evaluate your AI’s management quality before deployment.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Corvus ISR Reveals Synthetic Benchmark Results for Tracking Models

AIThis post was created with the assistance of artificial intelligence (AI).The published…

Chandigarh University Surges In Global Coverage

Chandigarh University has seen a surge in international media coverage, with GDELT reporting 14 mentions in a recent window, eight times higher than usual.

2026-08-24 – Interview – Interview With Petra Tschudin In The FuW

Swiss National Bank’s Petra Tschudin shares insights on monetary policy and financial stability during an interview with FuW, August 24, 2026.

Tornado Gluckstadt Windhose

A confirmed tornado struck Gluckstadt, Schleswig-Holstein, causing damage and evacuations. Details are still emerging about the extent and impact.