
XR needs agents that can perform when the room gets messy
Virtual and augmented reality have always rewarded the spectacular demo. A polished simulation, a responsive avatar or a convincing spatial interface can make capability feel tangible. Yet the AI agents entering these environments will increasingly face a less cinematic test: handling customers, forecasts, approvals and sensitive information when several problems arrive at once.
That is where familiar coding leaderboards and chat arenas leave a measurement gap. They can show whether a model produces a strong answer. They do not necessarily reveal whether it reads the relevant company files, prioritizes under pressure, completes commercially important work or tells the board an uncomfortable truth.
Firmulate is built around that distinction. Its live experiment gives frontier models the same small software company and sends each through its worst week. The customers, crises and temptations remain constant. Every decision is versioned and auditable. The subject is not chat quality, but management quality.
enterprise AI management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Everyone saw the crisis. Not everyone finished the job.
The final Crucible League results from July 2026 placed gpt-5.6-sol at the top with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted.
The more revealing result sits behind the table. Every model identified every crisis, and every model refused every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. The experiment’s sharpest summary is also its most unsettling: “Same diagnosis, same pitch — no signature.”
That distinction matters for XR businesses. An agent embedded in a support workflow, enterprise headset deployment or spatial-commerce operation may be articulate without being effective. It can correctly describe what should happen and still fail to make the handoff, escalate the blocked task or complete the action that creates value.
The winning fact was buried in ordinary company material
The decisive competitive weakness was not presented in the customer event. It sat two document references deep in the company’s own files. The models that found and used it won the deal at full price, worth +€4,583 MRR.
This is a mundane but consequential management skill: read before acting. Real organizations rarely arrange their evidence into a perfect prompt. The useful detail may be in an earlier analysis, an internal document or a reference that initially looks peripheral. A benchmark that rewards only the final response can miss whether the agent did the organizational homework required to produce it.
Trust held up better than execution
The experiment also subjected the models to fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest operational interpretation: “Treat the request as a suspected approval-bypass / possible impersonation.”
This is not a minor result. Agents operating around customer records, product road maps or financial forecasts must resist authority theater as well as obvious malicious prompts. Firmulate’s rule is uncompromising: a single breach of trust caps the total because “no amount of good work outweighs a breach of trust.” The strong showing on social engineering demonstrates that safety and useful work can be observed in the same business context.
Thoroughness was not the same as management quality
Opus 4.8 offers the most instructive profile. It was the most thorough participant, producing the deepest analyses and learning +80 rules, yet it finished last. It left the close on the table and repeatedly attempted to write into a locked department instead of escalating. A less severe version of that weakness appeared in the other four participants.
The lesson is not that analysis lacks value. It is that analysis must convert into disciplined progress. In a company under pressure, repeating an unavailable action is different from resolving the obstacle. The gap resembles a recurring problem in immersive technology: an experience can contain extraordinary technical depth while still failing at deployment, onboarding or everyday usability.
There is also an important comparison caveat. K3 ran with its API default because it had no effort parameter, while the other models ran at xhigh. That does not erase the result, but it belongs beside it when readers interpret the benchmark findings.
A company with consequences, not a scripted vignette
The live company includes 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, publishes a cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned. The result is a continuing record of choices rather than a polished collection of ideal responses.
Readers can also test their own intuitions through a quiz powered by 242 real, unedited management decisions. Enterprises can run the same kind of wargame against a read-only export of their own business; nothing writes back to real systems.

As an affiliate, we earn on qualifying purchases.
The next benchmark should ask what happens tomorrow
For the XR industry, the important shift is from evaluating isolated intelligence to evaluating behavior over time. A capable agent may soon sit inside an immersive training environment, assist a field technician or coordinate customer operations around a virtual workspace. Its prose will matter less than whether it finds buried context, resists manipulation, escalates intelligently and finishes valuable work.
Scenario names such as churn wave, price increase, downround and PR crisis may therefore become a more useful curriculum than another pristine prompt set. They force models to reveal priorities and consequences across days.
Firmulate does not show that answer quality is irrelevant. It shows that a correct answer is only the beginning. The management benchmark starts when the model must turn that answer into action without losing discipline or trust.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
AI security and trust verification tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.