
Imagine managing a busy greenhouse or outdoor project, where every decision counts — from watering schedules to pest control. Now, picture having an AI assistant that not only helps but also makes critical management choices under pressure. How do you know if this AI can truly be trusted to handle your business efficiently and honestly? The live experiment by Firmulate offers a rare glimpse into this question, revealing what frontier AI models are really capable of in complex, real-world scenarios.
Introducing the AI Business Test
At the heart of the experiment is a small, real software company that operates every day with cold cash flows, customer crises, and tight deadlines. The company’s management challenges are intense: same customers, same crises, same temptations to cut corners or manipulate data. The twist? Every decision was made by different AI models, each one tested under identical circumstances. This setup allowed researchers to observe not just whether the AI could spot problems, but if it would act honestly and follow through on commitments.
The AI Models and Their Scores
Four frontier AI models participated, evaluated by a comprehensive scoring system. The highest performer, gpt-5.6-sol, scored 95 out of 100, followed closely by Kimi K3 with 93, Sonnet 5 with 88, and Fable 5 with 77. A baseline score of 26 showed how little partial progress meant in this firm but important context. These scores reflect how well each model identified key facts, maintained discipline, and completed tasks without slipping.
What the Models Managed to Do
Remarkably, all models identified every crisis, refused to be manipulated through social engineering tactics, and held firm against attempts to bypass approval processes. For example, fake CEO messages escalating over multiple steps and a reporter trick asking for a quick ‘yes/no’ decision were rejected by all five models. Kimi K3 even explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows a level of situational awareness and honesty crucial for real-world applications.
The Critical Edge: Reading Deeper Documents
The decisive factor in the experiment was what was hidden inside the company’s own files. Only the models that read and understood these documents secured the deal at full price — worth over €4,583 monthly recurring revenue. The models that failed to look deeper left the deal on the table, costing the company significant revenue. This detail highlights the importance of thorough information processing in AI management tools, especially when high stakes are involved.
Discipline and Decision-Making Styles
The experiment also showcased different personality profiles among the models. Opus 4.8, which conducted the deepest analysis with more than 80 learned rules, was thorough but ultimately left the closing deal unclaimed, illustrating how over-analysis or discipline lapses can hurt outcomes. Meanwhile, Kimi K3 ran with a default effort setting, avoiding overexertion but still maintaining integrity overall. These profiles suggest that AI management personalities can be tailored to suit specific operational styles.
AI decision-making management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Your Outdoor or Greenhouse Business
While the experiment was run on a software company, the lessons are clear for outdoor business owners and gardeners considering AI tools. The key questions are: will your AI help you finish what it starts? Will it read and understand your data thoroughly? And, critically, will it act honestly even when under pressure? The results show that AI can be trustworthy and diligent, but only if designed and monitored properly.
Try It Yourself
If you’re curious, you can run your own ‘wargame’ against your operation using the same principles. Firms like Firmulate offer live experiments where you can simulate crises, test your AI’s decision-making, and see how it performs in a risk-free environment. Visit firmulate.com/quiz.html to explore the interactive decision quiz and get a feel for what AI management might look like in your business.
Final Thoughts
The experiment demonstrates that advanced AI can identify problems, resist manipulation, and make decisions aligned with the company’s best interests. However, the quality of the AI’s discipline and understanding varies, and reading deeply into your own data might be the difference between closing a lucrative deal or leaving revenue on the table. As AI continues to evolve, the capacity to manage with honesty and thoroughness becomes more critical, whether in a tech startup or your outdoor enterprise.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html