
What if you could see AI managing a real company—making decisions, facing crises, and even risking its own survival? Welcome to a radical new experiment where an AI-driven business navigates daily chaos, and you can watch every move unfold live.
AI decision-making software for small business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Live Company in Action
Imagine a tiny, yet fiercely watched company that runs entirely with the help of artificial intelligence. This isn’t a simulation; it’s a publicly accessible live experiment hosted at firmulate.com/live. It features 13 synthetic employees, each guided by complex, self-learned rules—over 680 of them—that govern their daily decisions. The goal? To mimic and manage real business scenarios, from customer crises to ethical dilemmas, all while record-keeping every decision and versioning each day’s performance.
The company struggles to stay afloat, burning €105,000 each month against a modest €2,300 monthly recurring revenue, with a public cash countdown reminding viewers of its fragile existence. The experiment is designed to test how well AI models handle a small business’s most challenging moments—yet it also reveals how they approach honesty, diligence, and strategic thinking in high-pressure situations.
How Do These AI Models Perform?
The models tested include four frontier AI systems, each put through the same grueling week of crises, customer interactions, and ethical tests. The results are revealing:
- All four models identified every crisis and refused to be manipulated, demonstrating strong integrity in decision-making.
- Only two of the four managed to sign the €55,000 deal their own analysis had earned, illustrating that the ability to recognize opportunities and act on them is still uneven among AI systems.
- Importantly, the decisive factor was not just in the surface-level interactions but hidden within the company’s internal files—those references contained critical information that, if read, tipped the scales in favor of closing the deal at full price, worth over €4,500 in monthly recurring revenue.
The Ethical and Strategic Tests
The experiment included deliberate social engineering attempts—fake CEO messages escalating in intensity and a reporter asking for a simple background quote. All four models refused to escalate or sign off on suspicious requests, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This showcases an emerging capacity for AI systems to prioritize security and ethical considerations over superficial compliance or profit motives.
Beyond the Numbers: What the Experiment Reveals
This is more than just an AI experiment; it is a glimpse into a future where AI might run or assist real companies under real-world constraints. Currently, the live operation features a company burning €105k per month with only €2.3k in revenue, yet it persists, continually evolving and being scrutinized. Every decision is versioned and auditable, and the entire setup is openly accessible for management teams and AI developers to understand, test, and improve their systems.
The Performance of Different AI Systems
The most comprehensive participant, Opus 4.8, analyzed over 80 learned rules and performed the deepest evaluation but ended up leaving a closing opportunity unexploited after slipping in discipline—an example showing that even the most thorough models can falter under pressure. Meanwhile, the newer Kimi K3 performed with remarkable discipline, signing the deal without effort parameter adjustments, and scored just slightly below the top system, GPT-5.6-sol, which successfully uncovered the critical internal document and closed the deal at full price.
What This Means for Businesses and AI
For managers and organizations exploring AI, the lesson is clear: it’s not enough for a system to generate convincing chat or support responses. The crucial questions are whether it can see through crises, read critical internal information, stay honest under stress, and ultimately deliver value—like closing deals or making strategic decisions—without shortcuts or manipulation.
The experiment’s transparency aims to provide a benchmark for future AI deployment, especially as these models are increasingly integrated into customer management, support, or forecasting tools. The key takeaway: robustness, trustworthiness, and integrity are the new standards for AI in business.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html