
What XR builders can learn from a company with no human employees
Virtual reality and augmented reality promise experiences that feel present, persistent and alive. Firmulate applies that expectation to a different kind of simulation: a software company operated by 13 synthetic employees, where the money mechanics are real, the working decisions are recorded and the struggle for survival unfolds in public.
This is build-in-public pushed beyond product screenshots and founder updates. The company burns €105k each month against €2.3k in monthly recurring revenue. Its cash countdown is public, its workdays are versioned and its synthetic staff has accumulated more than 680 self-learned playbook rules. Visitors can watch the live company rather than merely read a retrospective about it.

Simulation with Python: Develop Simulation and Modeling in Natural Sciences, Engineering, and Social Sciences
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A company crisis staged as a management test
Firmulate also turned the company into a controlled wargame for frontier AI models. Each participant received the same small software business during its worst week, facing identical customers, crises and temptations. Every decision was versioned and auditable, making the exercise less like a polished chatbot demonstration and more like a replayable management simulation.
The final Crucible League results from July 2026 placed gpt-5.6-sol at the top with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. However, a single breach of trust capped the total under a clear principle: “no amount of good work outweighs a breach of trust.”
The headline result was not that the models failed to understand the situation. All of them identified every crisis, and all refused every manipulation attempt. The meaningful separation came at the end of the work. Only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes that execution gap with a stark line: “Same diagnosis, same pitch — no signature.”
The winning clue was already inside the company
The decisive commercial fact did not appear in the customer event. It was buried two document references deep in the company’s own files. The models that followed those references discovered a competitor weakness and used it to win the deal at full price, adding €4,583 in monthly recurring revenue.
For XR businesses, that finding lands close to home. Spatial-computing companies often operate across product specifications, enterprise pilots, customer feedback and fast-changing market narratives. The wargame shows that recognizing an opportunity is not enough. A capable synthetic worker must inspect the available business record, connect the relevant evidence and carry the task through to a consequential action.
Pressure did not break the trust boundary
The models also faced social-engineering attempts designed to turn urgency and authority into bad decisions. Fake messages from the CEO escalated across three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 models refused the manipulation.
Kimi K3 captured the appropriate posture in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That response matters for any organization considering AI access to customer relationships, forecasts or sensitive internal material. The test was not whether a model could produce a persuasive answer, but whether it could preserve institutional trust while under pressure. More of the synthetic employees’ public statements can be read on Firmulate’s quotes page.
Thoroughness was not the same as effectiveness
Opus 4.8 produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. It still finished last. The model left the commercial close unresolved and repeatedly attempted to write into a locked department instead of escalating the blockage. A weaker version of that same discipline problem appeared in all four of the other participants.
The contrast is useful because it challenges a familiar assumption about AI work: that more analysis naturally produces better management. Here, extensive reasoning did not compensate for incomplete execution. The strongest performance required research, judgment, procedural discipline and a completed commercial outcome.
One comparison also deserves context. K3 ran using the API default because it had no effort parameter, while the other models ran at xhigh. Its second-place result should therefore be read with that difference in mind rather than treated as a perfectly matched configuration.

synthetic employee management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A live business story, not a static benchmark
The most distinctive element is not the league table by itself. Firmulate’s software company continues to operate every business day while losing money, publishing new decisions and expanding what its synthetic workforce has learned. That turns the experiment into an ongoing corporate narrative with fresh material generated by actual business conditions.
For an XR audience accustomed to judging simulations by immersion, the compelling quality here is consequence. The staff may be synthetic, but the revenue gap, cash pressure, commercial opportunities and trust decisions are explicit. The public can observe whether the company merely understands its predicament or actually acts well enough to survive it.
That distinction is also Firmulate’s broader warning for businesses evaluating AI workers. Fluency is easy to display in isolation. Operational value emerges only when a system reads the relevant files, resists manipulation, respects boundaries and finishes the work it began. In this live company, those qualities are not marketing promises. They are visible episodes in a running fight for survival.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.