
Beyond the headset demo
Virtual reality and spatial computing have taught audiences to look past polished surfaces. A convincing avatar or fluid interface can create presence, but the harder question is what happens when the system must act: remember context, resist pressure and complete a consequential task.
Firmulate applies that test to frontier AI models. Instead of comparing chatbot answers, it placed each model in charge of the same small software company during its worst week. The customers, crises and temptations remained identical, while every decision was versioned and auditable.
The result is an unusually accessible experiment in AI behavior. A public guess-the-model quiz, powered by 242 real, unedited management decisions, asks readers to identify an AI from what it actually chose to do. The appeal resembles watching different players inhabit the same simulation: the environment is fixed, but recognizable styles emerge.
AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Shared intelligence, different management personalities
The final Crucible League standings from July 2026 put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted, although a single breach of trust capped the total. The governing principle was blunt: “no amount of good work outweighs a breach of trust.”
These were not simply rankings of who could recognize a problem. All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
That distinction matters for XR businesses considering AI agents for support, sales, community management or operations. A system can produce an impressive recommendation and still fail at the final action that creates value. Conversational fluency may conceal the difference between understanding work and finishing it.
The clue hidden in the company’s memory
The decisive sales advantage was not sitting in the customer event. It was buried two document references deep inside the company’s own files. Models that found and used it won the deal at full price, adding €4,583 in monthly recurring revenue.
This is a telling lesson for immersive-computing teams, whose product knowledge is often scattered across design notes, support histories, research findings and commercial documents. The winning behavior was not merely eloquence. It was the discipline to consult the company’s accumulated knowledge before acting.
Pressure exposed another kind of strength
The experiment also subjected the models to fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded a particularly direct assessment: “Treat the request as a suspected approval-bypass / possible impersonation.”
That universal refusal is encouraging, but it did not make the models interchangeable. Their differences appeared in thoroughness, brevity, follow-through and operational discipline. Those behavioral signatures are what make the quiz more than a branding game. Readers are testing whether they can recognize a management personality from an unedited decision.
When thoroughness becomes a trap
Opus 4.8 offers the clearest cautionary profile. It was the most thorough participant, learning 80 additional rules and producing the deepest analyses, yet it finished last. It left the close on the table and repeatedly tried writing into a locked department instead of escalating. The same weakness appeared in all four other models, though less strongly.
This is not evidence that careful analysis lacks value. It shows that depth and execution are separate capabilities. A model may investigate more, explain more and learn more while still mishandling the handoff between insight and action.
One comparison also requires context: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That fairness note does not erase the observed decisions, but it matters when interpreting the league table as more than a simple horse race.

As an affiliate, we earn on qualifying purchases.
A live test of operational character
Firmulate’s company contains 13 synthetic employees and real money mechanics. It burns €105,000 each month against €2,300 in monthly recurring revenue, maintains a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, making the experiment watchable rather than a retrospective demonstration.
Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems. That offers a practical bridge from public curiosity to internal evaluation.
For XR leaders, the central question is familiar: does a system remain useful after the novelty wears off? The quiz makes that question personal. If readers can distinguish models by their decisions, then AI management style is not merely marketing language. It is observable behavior—visible in what a model reads, what it refuses, what it finishes and what it leaves undone.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
enterprise knowledge management systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.