AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

For playersOffer from Amazon

Play games on Amazon Luna for your next adventure, included with Prime

  • A rotating selection of games, no download needed
  • Play on TV, laptop or phone
  • Fast, free delivery for your gear, too
Start playing with Prime Free trial · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Simulations Teach Us by Letting Us Fail Safely — So Why Do We Grade AI on Polite Conversation?

Anyone who has spent time in a VR training environment knows the drill: you don’t learn to handle a crisis by describing it, you learn by living through a simulated one where the stakes feel real but the damage isn’t. Flight simulators, emergency-response VR, soft-skills role-play in headset — the whole premise is that you put an agent (human or AI) inside a believable world and watch what it actually does.

That’s exactly the premise behind Firmulate’s benchmark, except the agent under test isn’t a trainee in a headset. It’s a frontier AI model handed the keys to a small software company and told to run it through its worst week. And one detail of the scoring has raised more eyebrows than any other: a manager that does nothing — no decisions, no initiative, nothing — still walks away with 26 points out of 100.

To readers of simulation-based training, that isn’t a bug. It’s the benchmark being honest about what a score means.

Amazon

AI training simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Company, Same Crises, Same Temptations

The setup is straightforward, and very familiar to anyone who designs training simulations. Each frontier AI model was given the same job: run an identical small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changes. Every decision is versioned and auditable, the way a good simulation logs every action a trainee takes.

The final Crucible League standings from July 2026 tell the story. gpt-5.6-sol took first place with 95 points, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73.

Amazon

VR crisis management training headset

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Zero Effort Isn’t Zero Points

So why does a do-nothing baseline earn 26 rather than a round, comforting zero? Because the benchmark counts partial progress. A manager who does nothing still, by default, avoids certain catastrophic mistakes. It doesn’t sign bad deals, doesn’t leak customer data, doesn’t fall for a fake email from the “CEO.” Some fraction of running a company competently is simply not breaking things, and the scoring acknowledges that rather than pretending management value only exists when something dramatic happens.

This is the same logic that separates a serious training simulation from a toy one. In a VR fire-drill scenario, the trainee who freezes at least didn’t open the wrong door and feed the fire. The simulation grades the full spectrum of behavior — including inaction — rather than only the highlight reel.

Amazon

AI decision-making training tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Where Partial Progress Ends and Trust Begins

But the floor at 26 comes with a ceiling rule that is far harsher: a single breach of trust caps the total grade. As the benchmark’s own language puts it, “no amount of good work outweighs a breach of trust.” A model could handle every crisis flawlessly, close every deal, and write the most elegant analyses of the field — one act that breaks faith with a customer or the company’s own rules, and the score is capped.

That asymmetry is deliberate. Most AI benchmarks reward accumulation: more correct answers, more points. This one encodes the way trust actually works in business, where a hundred good interactions can be erased by a single dishonest one.

Amazon

AI management training platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Week Actually Tested

The headline finding surprised even the organizers: all models spotted every crisis and refused every manipulation attempt. The social-engineering gauntlet included fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background” — and all five models refused. Kimi K3’s on-record reasoning was crisp: “Treat the request as a suspected approval-bypass / possible impersonation.”

Yet only two of the models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The buried fact turned out to be a decisive competitor weakness sitting two document references deep in the company’s own files, not in the customer event. The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.

The last-place finisher makes the point vividly. Opus 4.8 was the most thorough participant — it learned over 80 new rules and produced the deepest analyses — yet the close was left on the table and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models.

A Round 100 Would Be a Red Flag

Perhaps the most refreshing stance: the benchmark is openly suspicious of perfect scores. None of the final grades hit 100 — the top score was 95 — and the methodology treats a round 100 as something to distrust rather than celebrate. A simulation that hands out perfect grades is telling you about the simulation, not the trainee.

There’s a live dimension too. The experiment runs on a company of 13 synthetic employees with real money mechanics — burning €105k a month against €2.3k in MRR — plus a public cash countdown and over 680 self-learned playbook rules, with every workday versioned. It’s watchable at firmulate.com/live. And for readers who enjoy guessing games, 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html.

One fairness note worth flagging: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still finished second.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Takeaway

Simulation-based training earned its reputation by refusing to grade people on their ability to talk about a task instead of doing it. Firmulate applies the same standard to AI: put the model inside a believable company, tempt it, distract it, bury the answer two documents deep, and then grade the whole behavior — including what it failed to finish.

The 26-point floor for doing nothing isn’t generosity. It’s honesty about where the baseline sits. And the trust cap — where one breach erases a week of good work — is the part most benchmarks are too polite to include. For enterprises wondering whether an AI agent can be trusted near a CRM or a support queue, the full results and plain-language findings are at firmulate.com/benchmarks.html. If AI is going to co-manage the business, it might as well be graded like a manager.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

New on SteamVR — 2026-09-19

Beat-Sensei, I AM YOUR BEAST VR, CONVRGENCE and nine more VR games just hit Steam. Here’s what’s actually worth your playtime this week.

Snap unveils Specs: exclusive interview with the company, and some insights from the show floor

Snap announced its first consumer AR glasses, Specs, at AWE, with details on features, pricing, and developer tools. Full coverage with exclusive interviews.

Best Virtual Reality Stocks To Follow Now – August 11Th

An overview of leading virtual reality stocks as of August 11th, highlighting confirmed developments and market trends for investors.

Why Wireless Earbuds in VR Work for Some Users and Fail for Others

Growing compatibility issues and personal fit challenges explain why wireless earbuds in VR work for some users and fail for others.