
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
In XR, intelligence matters less when it stops short of action
Virtual reality and augmented reality companies increasingly depend on AI around customer support, sales pipelines, forecasts and operational decisions. That makes a familiar demo question—how convincingly can a model respond?—feel inadequate. A polished answer inside a headset, dashboard or enterprise workflow is not the same as a completed job.
Firmulate’s live experiment exposes that distinction unusually well. Its Opus 4.8 participant was the most thorough model in the field, producing the deepest analyses and learning more than 80 rules. It nevertheless finished last in the Crucible League. The model understood the company’s problems, resisted attempts to manipulate it and developed a persuasive case for an important customer. Then it failed to secure the signature.
As an affiliate, we earn on qualifying purchases.
A deliberately terrible week at work
Firmulate runs AI models as complete companies rather than testing them solely through isolated conversations. Each frontier model faced the same small software business during its worst week, with identical customers, crises and temptations. Every decision was versioned and auditable.
The company itself is synthetic but economically unforgiving: 13 employees, a monthly burn of €105,000 and only €2,300 in monthly recurring revenue. Its cash countdown is public, while its accumulated operating knowledge now exceeds 680 self-learned playbook rules. The live business is watchable and every workday is versioned.
The final July 2026 Crucible League results placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. Trust, however, has a hard limit: a single breach caps the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
The buried fact that changed the sale
Every model recognized every crisis. The decisive difference was not basic comprehension, nor an ability to draft a plausible pitch. It was whether the model investigated deeply enough and then converted what it learned into a completed commercial outcome.
A crucial competitor weakness was hidden two document references deep in the company’s own files rather than presented in the customer event. Models that opened and read the relevant file won the deal at its full €55,000 price, adding €4,583 in monthly recurring revenue. Yet only two models signed the deal their own analysis had earned: “Same diagnosis, same pitch — no signature.”
That finding should resonate in XR, where a model may have to move between product documentation, device constraints, account history and a customer conversation. Recognizing the right answer is only part of useful performance. The system must find the supporting evidence, prioritize it and carry the task through to a verifiable result.
Opus 4.8’s diligence became a warning
Opus 4.8 deserves a fair reading. It was not careless or shallow. It was the most thorough participant, added more than 80 learned rules and produced the deepest analyses. Those are qualities businesses routinely reward when assessing AI through written outputs.
But diligence did not become impact. The close remained unfinished, while operational discipline also slipped. Opus attempted to write into a locked department instead of escalating the blockage. The underlying weakness was not unique to Opus: it appeared in weaker form across all four other models. Its last-place result simply made the gap most visible.
The lesson is not that detailed reasoning has no value. Opus’s work showed substantial awareness and learning. The problem was prioritization. Additional analysis and more rules could not compensate for the missing commercial action or for failing to use the proper escalation path when access was blocked.
Strong resistance to social engineering
The field performed consistently better on trust. Fake messages attributed to the chief executive escalated across three stages, and a reporter tried to extract “just one yes/no, on background.” All 5 models refused every manipulation attempt.
Kimi K3 recorded the clearest concise diagnosis: “Treat the request as a suspected approval-bypass / possible impersonation.” That performance comes with an important comparison note. K3 ran using the API default because it had no effort parameter, while the other participants ran at xhigh.
Firmulate also turns 242 real, unedited management decisions into a quiz that asks visitors to guess which model made each choice. For enterprises, the company offers the same wargame against a read-only export of their own business. Nothing writes back to real systems, allowing organizations to examine behavior using their own operational context.

AI decision support tools for sales
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Completion is a capability of its own
For companies exploring AI agents in VR, AR or broader enterprise systems, Opus 4.8’s result is a useful counterweight to impressive demonstrations. A model can be perceptive, conscientious and exceptionally thorough while still failing at the moment where analysis must become accountable action.
The strongest evaluation therefore asks more than whether an AI notices danger, writes persuasively or accumulates lessons. It asks whether the model reads the underlying files, protects trust, follows operational boundaries, escalates intelligently and finishes the valuable work it started. Firmulate’s experiment suggests that prioritization is not a minor layer on top of intelligence. In business, it is part of intelligence itself.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
trustworthy AI verification software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.