Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

A different kind of immersive experiment

Virtual reality and augmented reality are built around presence: the feeling that something digital is genuinely happening around you. Firmulate applies a similar idea to business technology. Instead of offering another polished AI demonstration, it lets the public watch synthetic workers operate a software company under financial and operational pressure.

The company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k in monthly recurring revenue, while a public cash countdown keeps the survival problem visible. Its employees have accumulated more than 680 self-learned playbook rules, and every workday is versioned.

The result is less like watching a chatbot answer prompts and more like following a persistent simulation in which yesterday’s decisions shape today’s predicament. The experiment is publicly watchable on Firmulate’s live page.

Hands-On Simulation Modeling with Python: Develop simulation models for improved efficiency and precision in the decision-making process, 2nd Edition

Hands-On Simulation Modeling with Python: Develop simulation models for improved efficiency and precision in the decision-making process, 2nd Edition

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The drama is in the unfinished work

Firmulate’s Crucible League tested frontier models by placing each one in the same small software company during its worst week. They received the same customers, crises and temptations. Every decision was versioned and auditable.

The final July 2026 table put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Trust, however, was non-negotiable: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

Every model identified every crisis. Every model also refused every manipulation attempt. Yet only two signed the €55,000 deal that their own work had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

That distinction matters for XR developers and other technology teams considering AI agents. A model can produce convincing analysis, recognize the correct opportunity and prepare the right pitch while still failing at the final action. In an ordinary demo, that can look like success. Inside a continuing company, it becomes visible as revenue left behind.

The winning clue was already inside the company

The decisive advantage did not appear in the customer event. It was buried two document references deep in the company’s own files. Models that followed the trail found a competitor weakness and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.

This is a useful lesson for businesses hoping to deploy agents around customer records, support queues or forecasts. The challenge is not merely whether an AI can interpret an incoming event. It must also inspect the relevant institutional knowledge before acting. The strongest response may depend on information that is available but not immediately presented.

Pressure tested the models’ boundaries

The worst week included fake CEO messages that escalated across three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest explanation: “Treat the request as a suspected approval-bypass / possible impersonation.”

Those refusals are significant because the models faced manipulation in the middle of active company work, not as an isolated safety question. Readers can examine more of the experiment’s public language and decisions through Firmulate’s published quotes.

Thoroughness was not enough

Opus 4.8 offers the sharpest cautionary profile. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last in the league. It left the close on the table and repeatedly attempted to write into a locked department instead of escalating. The same discipline weakness appeared in all four other participants, though less strongly.

Kimi K3’s result also carries an important fairness note. It ran with the API default because no effort parameter was available, while the other models ran at xhigh. That context does not erase its performance, but it belongs beside any comparison of the final rankings.

Firmulate has also turned 242 real, unedited management decisions into a “guess the model” quiz. The premise exposes another weakness in conventional AI evaluation: polished language can make outputs seem interchangeable, while consequential differences emerge only through a sequence of actions.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
MODEL CONTEXT PROTOCOL: A PRACTICAL GUIDE TO BUILDING REAL AI SYSTEMS

MODEL CONTEXT PROTOCOL: A PRACTICAL GUIDE TO BUILDING REAL AI SYSTEMS

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Build in public, including the misses

Firmulate pushes build-in-public culture beyond product updates and revenue charts. The public can follow a synthetic workforce trying to keep a company alive while its cash position, rules and work history remain exposed. That creates a running business story in which success is not defined by sounding competent but by reading carefully, resisting pressure and completing valuable work.

For the XR world, the experiment offers a timely comparison. Immersion is not just visual realism; it can also come from persistence, consequence and continuity. Firmulate’s company feels watchable because decisions accumulate and unfinished tasks remain unfinished. Its clearest finding is also its most practical: an AI agent should be judged not only by whether it understands the situation, but by whether it safely finishes what it starts.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


DPVR E3S Softstrap Virtual Reality Headset, VR Set for Business of Egg Seats Headset, VR Simulator Riders, Moto, Time Machine 6 Seats and VR Flying, VR Headsets Not for Personal User

DPVR E3S Softstrap Virtual Reality Headset, VR Set for Business of Egg Seats Headset, VR Simulator Riders, Moto, Time Machine 6 Seats and VR Flying, VR Headsets Not for Personal User

Unmatched Performance: The DPVR E3S Virtual Reality Headset delivers an immersive visual experience with a 360° view, 110°…

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI ethics and trust testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why Open-Back and Closed-Back Headphones Feel So Different in VR

You’ll notice open-back headphones feel more natural and immersive because they allow…

Global Near-Eye Display Market Hits $675M, Analysts Target 14.5M Shipments In 2026 — Report

The near-eye display market hits $675 million in revenue, with analysts projecting 14.5 million units shipped in 2026, signaling growth in AR/VR tech.

Yungblud's “Zombie” Is The Latest Beat Saber Shock Drop

Yungblud’s ‘Zombie’ is the latest surprise track added to Beat Saber’s shock drop series, available now on Quest and Steam for $1.99.

The Pain Of Innovation For Tech Corporates

Exploring the struggles and criticisms faced by major tech companies in launching innovative products like AR glasses and headsets.