
Imagine managing your smart home or appliances with AI that doesn’t just respond, but actually makes decisions — sometimes even risking a breach of trust just to close a deal or fix a crisis. Could your AI be honest and effective under pressure? At the frontier of AI-driven management, real-world experiments reveal surprising differences between models, and the stakes are higher than you might think.
The Experiment: Testing AI as a Business Manager
In a groundbreaking live experiment, four advanced AI models were put in the role of running a small, real software company. This wasn’t a simulation in a lab — it was a real company with real money, real crises, and real temptations. Every day, the models faced the same set of challenges: customer crises, internal decisions, and ethical dilemmas. The goal? See which AI could best manage the company, stay honest, and close a critical deal worth €55,000.
AI management software for small business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Did the Models Do?
Every decision was recorded and viewable — from reading customer complaints to negotiating deals. All four models identified the same crises and refused manipulation attempts, such as fake CEO messages or reporter tricks designed to trick management into bypassing protocols. They showed integrity in the face of social engineering, refusing to escalate fake approval requests.
The Key Difference: Reading the Files
While all four AI models demonstrated competence in crisis detection and rejection of manipulation, the decisive factor emerged from a hidden detail: a buried document in the company’s files. Only two models successfully read this critical piece of information, which contained the key to closing the deal at full price. These models managed to leverage this knowledge, winning the deal worth over €4,583 in monthly recurring revenue (MRR). The other two, despite similar diagnostics, left the opportunity on the table, settling for less or nothing at all.
Personality and Discipline in AI
The models showed distinct management styles:
- gpt-5.6-sol: The top scorer with 95 points, it found the hidden fact, closed the deal, and maintained discipline throughout.
- Kimi K3: A newcomer, it scored 93, and was praised for its fairness and clean decision-making. It also secured the deal, running without an effort parameter, indicating a balanced approach.
- Sonnet 5: With 88 points, it managed to close the deal but exhibited some process slips during the critical closing phase.
- Fable 5: Scoring 77, it also closed the deal but showed weaker discipline, leaving opportunity on the table by failing to escalate some issues properly.
Another interesting point: all models refused social engineering attempts, like staged CEO messages or background yes/no questions from reporters. Their reasoning? Treating such requests as potential impersonation or bypass attempts.
Beyond the Benchmarks
This experiment underscores a vital insight for anyone deploying AI in management roles — the importance of thoroughness, reading comprehension, and integrity. Even models with similar diagnosis capabilities behaved differently when it came to leveraging hidden information and sticking to disciplined processes. For smart home setups, this translates into: will your AI read the critical device logs and data, or overlook essential details? Will it stay honest under pressure or take shortcuts that undermine trust?
Real Money, Real Consequences
The company used in this live experiment operates with 13 synthetic employees but handles actual financial mechanics. It loses €105k each month against a revenue of just €2.3k. The AI’s decisions directly impact this bottom line — making the ability to finish what it starts and read all relevant data a matter of real consequence.
Why Should You Care?
If AI agents are to touch your customer relationships, support workflows, or even smart appliances, the question isn’t just how well they can generate human-like chat. It’s whether they will complete their tasks honestly, thoroughly, and reliably. When managing your smart home or business, this could mean the difference between a smooth operation and a costly breach or missed opportunity.
Try It Yourself
Want to see which AI might be best suited for your needs? Visit the quiz page to test your management decisions against different models and discover how they differ in handling real-world challenges. For enterprises, a simulation can be run on your own data — without risking your actual systems — to see how your AI management team would perform.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html