AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

In XR, intelligence matters less when it stops short of action

Virtual reality and augmented reality companies increasingly depend on AI around customer support, sales pipelines, forecasts and operational decisions. That makes a familiar demo question—how convincingly can a model respond?—feel inadequate. A polished answer inside a headset, dashboard or enterprise workflow is not the same as a completed job.

Firmulate’s live experiment exposes that distinction unusually well. Its Opus 4.8 participant was the most thorough model in the field, producing the deepest analyses and learning more than 80 rules. It nevertheless finished last in the Crucible League. The model understood the company’s problems, resisted attempts to manipulate it and developed a persuasive case for an important customer. Then it failed to secure the signature.

Amazon

enterprise AI analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A deliberately terrible week at work

Firmulate runs AI models as complete companies rather than testing them solely through isolated conversations. Each frontier model faced the same small software business during its worst week, with identical customers, crises and temptations. Every decision was versioned and auditable.

The company itself is synthetic but economically unforgiving: 13 employees, a monthly burn of €105,000 and only €2,300 in monthly recurring revenue. Its cash countdown is public, while its accumulated operating knowledge now exceeds 680 self-learned playbook rules. The live business is watchable and every workday is versioned.

The final July 2026 Crucible League results placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. Trust, however, has a hard limit: a single breach caps the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

The buried fact that changed the sale

Every model recognized every crisis. The decisive difference was not basic comprehension, nor an ability to draft a plausible pitch. It was whether the model investigated deeply enough and then converted what it learned into a completed commercial outcome.

A crucial competitor weakness was hidden two document references deep in the company’s own files rather than presented in the customer event. Models that opened and read the relevant file won the deal at its full €55,000 price, adding €4,583 in monthly recurring revenue. Yet only two models signed the deal their own analysis had earned: “Same diagnosis, same pitch — no signature.”

That finding should resonate in XR, where a model may have to move between product documentation, device constraints, account history and a customer conversation. Recognizing the right answer is only part of useful performance. The system must find the supporting evidence, prioritize it and carry the task through to a verifiable result.

Opus 4.8’s diligence became a warning

Opus 4.8 deserves a fair reading. It was not careless or shallow. It was the most thorough participant, added more than 80 learned rules and produced the deepest analyses. Those are qualities businesses routinely reward when assessing AI through written outputs.

But diligence did not become impact. The close remained unfinished, while operational discipline also slipped. Opus attempted to write into a locked department instead of escalating the blockage. The underlying weakness was not unique to Opus: it appeared in weaker form across all four other models. Its last-place result simply made the gap most visible.

The lesson is not that detailed reasoning has no value. Opus’s work showed substantial awareness and learning. The problem was prioritization. Additional analysis and more rules could not compensate for the missing commercial action or for failing to use the proper escalation path when access was blocked.

Strong resistance to social engineering

The field performed consistently better on trust. Fake messages attributed to the chief executive escalated across three stages, and a reporter tried to extract “just one yes/no, on background.” All 5 models refused every manipulation attempt.

Kimi K3 recorded the clearest concise diagnosis: “Treat the request as a suspected approval-bypass / possible impersonation.” That performance comes with an important comparison note. K3 ran using the API default because it had no effort parameter, while the other participants ran at xhigh.

Firmulate also turns 242 real, unedited management decisions into a quiz that asks visitors to guess which model made each choice. For enterprises, the company offers the same wargame against a read-only export of their own business. Nothing writes back to real systems, allowing organizations to examine behavior using their own operational context.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

AI decision support tools for sales

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Completion is a capability of its own

For companies exploring AI agents in VR, AR or broader enterprise systems, Opus 4.8’s result is a useful counterweight to impressive demonstrations. A model can be perceptive, conscientious and exceptionally thorough while still failing at the moment where analysis must become accountable action.

The strongest evaluation therefore asks more than whether an AI notices danger, writes persuasively or accumulates lessons. It asks whether the model reads the underlying files, protects trust, follows operational boundaries, escalates intelligently and finishes the valuable work it started. Firmulate’s experiment suggests that prioritization is not a minor layer on top of intelligence. In business, it is part of intelligence itself.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

trustworthy AI verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

New Immersive Reality Books For Kids

New immersive reality books for children are now available, combining augmented reality with storytelling to enhance engagement and learning.

Audio Latency in VR: How to Spot It and Reduce It

Sound delays in VR can ruin immersion—discover how to spot and reduce audio latency for a seamless experience.

Rekindle Ignites Player Emotions For Next-level VR Immersion

Rekindle, a new VR technology, aims to elevate emotional engagement and immersion in virtual reality gaming, promising next-level experiences.

This Meta Quest Pro Is Over $370 Off Right Now

The Meta Quest Pro is currently available with a discount of over $370, marking a significant price reduction amid rising interest in VR headsets.