
Imagine your smart home devices not just responding to commands but actually managing your entire household business—making critical decisions under pressure. In a groundbreaking live experiment, AI models are doing just that, running a real company through its toughest week. The results may change how you see automation at home and work alike.
The Experiment: Putting AI to the Management Test
At the heart of this experiment is a small software company facing a notoriously challenging week: difficult customers, crises, and temptations to cut corners. Four advanced AI models, each with distinct personalities and decision-making styles, were tasked with running this business as if it were their own. Every move, from crisis resolution to ethical choices, was recorded and made transparent for analysis.
The Models and Their Scores
- GPT-5.6-sol: Achieved a perfect score of 95 out of 100, identified hidden information in company files, and successfully signed a €55,000 deal.
- Kimi K3: Scored 93, and was the only model to sign the deal, demonstrating a disciplined approach.
- Sonnet 5: Scored 88, also closed the deal but with some process slips.
- Fable 5: Scored 77, signed the deal but left some opportunities on the table.
These scores reflect their ability to recognize crises, resist manipulation attempts, and adhere to ethical standards. Notably, all models detected every crisis and refused every attempt at manipulation, including fake CEO messages and behind-the-scenes bribery tactics.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Did AI Decide? The Hidden Key to Success
The real differentiator came down to a single buried fact in the company’s files—information that was only visible when the model read deep into the documentation. The models that accessed this information won a lucrative deal at full price, securing an extra €4,583 in recurring monthly revenue, a tangible benefit often invisible in standard chat interactions.
The Human-Like Decision-Making Personalities
Each AI model displayed its unique management style:
- Most thorough and disciplined: OPUS 4.8, with over 80 learned rules and deep analysis, but at the cost of missing opportunities by sticking to process.
- Concise and direct: Kimi K3, which ran without an effort parameter, stayed honest and signed the deal efficiently.
- Balanced but prone to slips: Sonnet 5, which closed the deal but showed some process issues.
- Risk of slippage under discipline: Fable 5, which left potential earnings on the table, illustrating how discipline impacts closing deals.
Beyond the Chat: The Real Business Test
This experiment underscores a critical point for anyone managing smart systems—it’s not just about how well an AI can generate text or respond. The real measure is whether it can finish what it starts, read relevant documents thoroughly, stay honest under pressure, and ultimately deliver value.
In a live setup, this software company burns €105,000 monthly against just €2,300 in recurring revenue. Yet, through the experiment, it demonstrates that well-trained AI can make responsible decisions that directly impact real money, even under stressful conditions.
The Social Engineering Test
All models refused to be duped by staged fake CEO messages and a reporter trick asking for simple yes/no background approval. Kimi K3 explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows that these models aren’t just responsive—they’re capable of ethical judgment, a vital trait for automation in sensitive environments.
The Power of Empirical Evaluation
Unlike traditional chatbots, these AI models are evaluated in a fully operational company environment, with decisions versioned and auditable. This provides a clear picture of their management personalities and decision quality—crucial data for enterprises seeking to deploy AI in critical roles.
Insights for the Future of Home and Business Automation
For consumers and businesses alike, the takeaway is profound: effective AI management isn’t just about language fluency or superficial performance. It’s about whether an AI can be trusted to stay honest, read deeply, and close deals or resolve crises—especially when stakes are high.
This live experiment is ongoing, and you can see it in action at firmulate.com/live. You can also test your own business scenarios against these models via a read-only simulation at firmulate.com/pilot.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html