
In a world obsessed with AI scores and performance metrics, it’s easy to think that the highest scores mean the most capable models. But what if the baseline—a do-nothing AI—still scores 26 out of 100? That’s the surprising insight from a recent public benchmark experiment, and it raises important questions about how we evaluate AI systems for real-world business tasks.
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
Understanding the Baseline and What It Tells Us
In the latest iteration of the Crucible League, four frontier AI models were tested by running a simulated small software company through its worst week. This wasn’t a scripted chat demo; it was a live, auditable test where each model faced the same crises, customer demands, and temptations to cheat. The goal: measure management quality, not just language prowess.
Remarkably, even the so-called ‘do-nothing’ baseline—a simple run with no intelligent intervention—scored 26 points out of 100. Why? Because partial progress counts. For example, merely reading critical documents to find a buried fact, or refusing manipulation attempts, adds to that score. This demonstrates that an honest benchmark values tangible decision-making and ethical behavior, not just clever language generation.
As an affiliate, we earn on qualifying purchases.
Why a Zero Score Isn’t Realistic
You might expect a baseline with no effort to score zero, but the experiment shows otherwise. A do-nothing approach still detects crises, refuses manipulation, and sometimes even finds crucial information hidden in files—contributing to that 26-point floor. It’s an acknowledgment that even minimal effort can yield some value in managing complex business scenarios.
AI ethics and trustworthiness tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Full Progress vs. Trust Breaches
Another key finding is that a single breach of trust can cap the overall score. All models performed well in crisis detection and manipulation refusal—but when one signed a fake deal without proper analysis, it capped the total performance. That’s because trustworthiness and consistency are weighted heavily; no amount of good work can outweigh a breach of integrity.
enterprise AI decision-making platforms
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real-World Stakes and Discrepancies
The live experiment included a simulation of a real company with 13 synthetic employees, handling actual money mechanics—burning €105,000 monthly against €2,300 in revenue, with every rule and decision versioned and auditable. The models demonstrated they could identify crises, refuse manipulative requests, and read files successfully. Yet, only two models signed the full-priced deal, illustrating the gap between raw technical capability and disciplined execution.
Model Performance Highlights
- gpt-5.6-sol scored 95 points, successfully finding a buried fact and closing the deal.
- Kimi K3 scored 93, also closing at full price, with the cleanest discipline of the group.
- Sonnet 5 scored 88, closing the deal with minor slips.
- Opus 4.8, with the most thorough analysis, scored 77, leaving the close on the table due to discipline lapses.
As an affiliate, we earn on qualifying purchases.
Implications for Business and AI Evaluation
This experiment reveals that the focus should be on management quality—reading files thoroughly, refusing manipulation, and staying disciplined—rather than just language fluency. It also underscores the importance of trustworthiness. A single breach caps the score, emphasizing that ethical behavior under pressure is critical for AI deployment in real-world settings.
What This Means for Your Organization
If your business relies on AI agents to handle customer relations, support, or forecasting, it’s not enough that they write well. The real question is whether they finish what they start, read your files carefully, stay honest, and deliver valuable work. Running simulations—like the one at Firmulate—can help you assess your AI’s management quality before deployment.
Get Started with Fair and Transparent Benchmarking
For enterprises interested in testing their AI systems, Firmulate offers live experiments where models are evaluated against real crises, money mechanics, and temptations—without ever risking actual systems. You can see how your AI measures up in a transparent, auditable environment. Visit firmulate.com/benchmarks.html to learn more about how your AI workforce performs in a realistic business simulation.

The first step in trusting AI for business is understanding its real capabilities under stress. A do-nothing baseline scores 26 points—showing that even minimal effort yields value. The true test is whether AI can finish what it starts, stay honest, and manage risks effectively. Firmulate’s live benchmarking offers a transparent way to evaluate your AI’s management quality before deployment.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
