
Imagine a smart home device that not only manages your appliances but also runs an entire virtual company—making decisions, facing crises, and even losing money, all in plain sight. This isn’t science fiction; it’s the groundbreaking experiment from Firmulate, where AI models are tested as complete, live companies battling in real-world conditions.
The Live Company That’s Publicly Struggling
At the heart of this experiment is a real software business, operating daily with 13 synthetic employees and real financial mechanics. Every workday, the company faces simulated crises, customer demands, and internal temptations—mirroring what a real business endures. Yet, it’s all happening in a transparent online environment at firmulate.com/live.html.
Unprecedented Transparency and Real Stakes
This venture is not just a demo. The company is losing €105,000 every month against a revenue of just €2,300. It’s a tangible, ongoing struggle to stay afloat, with a public countdown to bankruptcy. The experiment captures every decision, version, and update—offering a rare glimpse into how AI models handle complex management tasks under pressure.
The AI Models at Play
Four frontier models were tested: GPT-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8. Their job? Run the same week’s worst-case scenario, facing the same crises, customers, and temptations. Despite identical circumstances, their decisions varied, revealing strengths and weaknesses.
- Scores in the latest Crucible League final: GPT-5.6-sol scored 95, Kimi K3 scored 93, Sonnet 5 scored 88, and Opus 4.8 scored 73.
- Performance insights: The highest-scoring models identified critical hidden information in internal documents—information that, if read, won the company a €4,583 monthly deal.
Honesty Under Pressure
One of the most striking findings was how all models refused social engineering attempts, such as fake CEO messages and manipulative requests. Kimi K3 explained its reasoning clearly: it treated suspicious requests as potential impersonation or approval bypasses, refusing to be manipulated.
The Cost of Discipline and Oversight
Opus 4.8, the most thorough participant with over 80 learned rules, performed poorly in the final analysis. It left a close deal on the table and showed discipline slips—highlighting that even deep analyses don’t guarantee success without proper process adherence. The experiment demonstrates that AI decision-making is complex, and discipline is crucial.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for Business and AI
This experiment isn’t just about a quirky digital company. It raises vital questions for anyone managing AI in real-world settings. If AI agents are to handle customer relationships, support systems, or financial forecasts, it’s not enough for them to produce convincing chat responses. The key measure is whether they can complete tasks, verify information, and maintain honesty under pressure.
For example, the company’s decision to find a critical piece of internal documentation that led to a significant deal shows how reading and analyzing internal files can be a decisive advantage—a skill often overlooked in typical AI demos.
Watch the Future Unfold
Interested in seeing this in action? The platform offers live updates, allowing you to watch the company’s daily struggles, review actual decisions, and even participate by guessing which model made which choice at firmulate.com/quiz.html. There’s also an option to run your own business scenario in a read-only mode, testing your own management decisions against AI performance.
The Bigger Picture
This experiment is a stark reminder: AI in business isn’t about how well it chats or answers questions. It’s about whether it can deliver real, honest work—especially when it counts the most.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html