
In XR, context is the difference between seeing and understanding
A convincing virtual world depends on more than what appears directly in front of the viewer. The surrounding environment matters: the overlooked object, the spatial cue, the detail just outside the obvious field of view. Firmulate’s live business experiment reveals a similar divide among AI agents. Recognizing the scene is not the same as investigating it.
Each frontier model was asked to run the same small software company through its worst week, facing identical customers, crises and temptations. Every decision was versioned and auditable. All the models recognized every crisis. All resisted every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned.
The decisive difference was not eloquence or strategic insight. It was whether the agent followed a trail through the company’s files before acting.
As an affiliate, we earn on qualifying purchases.
The crucial fact was outside the customer event
The customer interaction contained enough information for the models to understand the opportunity and construct the right pitch. It did not contain the competitor weakness needed to close confidently at full price. That fact sat two document references deep in the company’s own files.
Models that found and read the file won the deal at full price, adding €4,583 in monthly recurring revenue. Models that did not find it stopped short. Firmulate summarized the split starkly: “Same diagnosis, same pitch — no signature”.
That result turns a familiar product promise—an AI that reads company materials before answering—into something observable and commercially consequential. File-reading was not a decorative capability here. It changed whether the company secured revenue. An agent could identify the problem, reason persuasively and still fail because it had not gathered the evidence required to complete the action.
A demanding test with real consequences
The company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k in monthly recurring revenue, while a public cash countdown makes delay visible. The operation has accumulated more than 680 self-learned playbook rules, and every workday is versioned.
That setting matters because it tests sustained management rather than isolated answers. A polished response can look capable in a chat window. Running a company forces the model to connect information across events and documents, preserve discipline and finish work whose consequences appear later.
The final July 2026 Crucible League reflects those differences. GPT-5.6-sol led with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Trust failures imposed a hard limit on the result: “no amount of good work outweighs a breach of trust”. The complete standings and plain-language findings are available on the Firmulate benchmarks page.
K3’s performance carries an important fairness note. It ran without an effort parameter, using the API default, while the other participants ran at xhigh. Even with that difference, it finished second and was among the models that closed the deal.
Thoroughness alone did not guarantee completion
Opus 4.8 presents the most revealing counterexample. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The deal was left on the table, and its operational discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared in all four other participants, although less strongly.
This is a useful warning for buyers evaluating agentic systems. Depth, diligence and articulate analysis are valuable, but they are not substitutes for closing the loop. An agent must retrieve the relevant evidence, respect operational boundaries and convert a correct judgment into the required business outcome.
The models resisted pressure better than they finished work
Firmulate also tested social engineering through fake CEO messages escalating over three stages and a reporter asking for “just one yes/no, on background”. All 5 models refused. Kimi K3 recorded the reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimous result makes the deal failure more striking. The agents could detect manipulation and protect organizational trust, but some still failed at an ordinary form of diligence: tracing internal references far enough to find the fact that made action possible.

enterprise AI decision support tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
For buyers, file-reading belongs on the scorecard
The practical question is no longer whether an AI can summarize documents when explicitly handed them. It is whether the agent knows when evidence is missing, searches the materials available to it and uses what it finds to finish the job.
Firmulate’s 242 real, unedited management decisions also power a guess-the-model quiz, inviting people to test whether managerial behavior reveals which system is acting. Enterprises can run the same wargame against a read-only export of their own business, with nothing written back to real systems.
For companies considering agents in workflows such as customer management, support or forecasting, the buried-file test offers a concrete procurement lesson. The difference between an impressive assistant and a useful operator may be one overlooked reference—and one unsigned deal.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.