AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Beyond the headset demo

Virtual reality and spatial computing have taught audiences to look past polished surfaces. A convincing avatar or fluid interface can create presence, but the harder question is what happens when the system must act: remember context, resist pressure and complete a consequential task.

Firmulate applies that test to frontier AI models. Instead of comparing chatbot answers, it placed each model in charge of the same small software company during its worst week. The customers, crises and temptations remained identical, while every decision was versioned and auditable.

The result is an unusually accessible experiment in AI behavior. A public guess-the-model quiz, powered by 242 real, unedited management decisions, asks readers to identify an AI from what it actually chose to do. The appeal resembles watching different players inhabit the same simulation: the environment is fixed, but recognizable styles emerge.

Amazon

AI decision-making software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Shared intelligence, different management personalities

The final Crucible League standings from July 2026 put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted, although a single breach of trust capped the total. The governing principle was blunt: “no amount of good work outweighs a breach of trust.”

These were not simply rankings of who could recognize a problem. All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

That distinction matters for XR businesses considering AI agents for support, sales, community management or operations. A system can produce an impressive recommendation and still fail at the final action that creates value. Conversational fluency may conceal the difference between understanding work and finishing it.

The clue hidden in the company’s memory

The decisive sales advantage was not sitting in the customer event. It was buried two document references deep inside the company’s own files. Models that found and used it won the deal at full price, adding €4,583 in monthly recurring revenue.

This is a telling lesson for immersive-computing teams, whose product knowledge is often scattered across design notes, support histories, research findings and commercial documents. The winning behavior was not merely eloquence. It was the discipline to consult the company’s accumulated knowledge before acting.

Pressure exposed another kind of strength

The experiment also subjected the models to fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded a particularly direct assessment: “Treat the request as a suspected approval-bypass / possible impersonation.”

That universal refusal is encouraging, but it did not make the models interchangeable. Their differences appeared in thoroughness, brevity, follow-through and operational discipline. Those behavioral signatures are what make the quiz more than a branding game. Readers are testing whether they can recognize a management personality from an unedited decision.

When thoroughness becomes a trap

Opus 4.8 offers the clearest cautionary profile. It was the most thorough participant, learning 80 additional rules and producing the deepest analyses, yet it finished last. It left the close on the table and repeatedly tried writing into a locked department instead of escalating. The same weakness appeared in all four other models, though less strongly.

This is not evidence that careful analysis lacks value. It shows that depth and execution are separate capabilities. A model may investigate more, explain more and learn more while still mishandling the handoff between insight and action.

One comparison also requires context: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That fairness note does not erase the observed decisions, but it matters when interpreting the league table as more than a simple horse race.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI management simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A live test of operational character

Firmulate’s company contains 13 synthetic employees and real money mechanics. It burns €105,000 each month against €2,300 in monthly recurring revenue, maintains a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, making the experiment watchable rather than a retrospective demonstration.

Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems. That offers a practical bridge from public curiosity to internal evaluation.

For XR leaders, the central question is familiar: does a system remain useful after the novelty wears off? The quiz makes that question personal. If readers can distinguish models by their decisions, then AI management style is not merely marketing language. It is observable behavior—visible in what a model reads, what it refuses, what it finishes and what it leaves undone.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

enterprise knowledge management systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI decision audit tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

XREAL Aura Hybrid AR and VR Glasses: Hands on Guide

An in-depth look at the XREAL Aura hybrid AR and VR glasses, including features, user experience, and what remains to be seen from the device.

A.R. Rahman Brings His Immersive Music Experiences To Apple Vision Pro

A.R. Rahman introduces ARR Immersive for Apple Vision Pro, offering 180- and 360-degree VR music and performance content. Content now available for download.

Bramblefort Aspires To Be Equal Parts Survival Horror And Immersive Sim

Bramblefort, a VR survival horror game inspired by Resident Evil 4 and Bloodborne, aims to blend horror and immersive sim elements. Full release pending.

XREAL Aura Hybrid AR And VR Glasses: Hands On Guide

A detailed guide to XREAL’s Aura Hybrid glasses, highlighting confirmed features, design, and what to expect from this mixed reality device.