
Imagine a garden where the strongest plants survive not by just growing tall but by resisting pests, drought, and disease. Now, picture managing a business with the same resilience. Recent experiments with artificial intelligence (AI) show that some models can outperform human decision-making in managing a company’s toughest week — refusing temptation, uncovering hidden risks, and sealing profitable deals. It’s a surprising story for anyone interested in sustainable growth, whether you’re tending a garden or growing a business.
Get garden gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Experiment: Simulating a Business Crisis
Four advanced AI models faced the same challenge: run a small software company through its worst week. This wasn’t a simple test of chat skills; it was a rigorous simulation involving real crises—customer churn, security breaches, and manipulative social engineering attempts. Every decision was carefully tracked and auditable, mimicking real management decision-making. The goal was to see which model could identify risks, resist temptation, and close profitable deals under pressure.
Results that Defy Expectations
All four models successfully identified every crisis and refused manipulation attempts. This demonstrates a remarkable capacity for honesty and discernment—key traits for managing real business risks. However, only two of the four models managed to close a €55,000 deal that their own analysis had earned them. The difference? One model found a buried document reference that revealed a crucial security flaw—allowing it to close the deal at full price, adding +€4,583 MRR.
The Surprise Winner: The Newcomer
The top scorer, gpt-5.6-sol, led with a score of 95 out of 100. Close behind was Kimi K3 from Moonshot, scoring 93. Despite being a newcomer, K3 demonstrated the cleanest discipline—resisting all baits and only deviating once in a disciplined manner. The other models, Sonnet 5 and Fable 5, scored 88 and 77 respectively, with Opus 4.8 trailing at 73. They all managed the crises but showed weaknesses in process discipline and risk detection.
AI business decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Edge: Reading Beyond the Surface
The crucial edge for K3 wasn’t just in crisis diagnosis but in reading deeper into company files. The experiment revealed that the decisive advantage lay two document references deep in internal files, not in external customer interactions. Models that read and analyze these internal documents were able to uncover risks and seize opportunities that others missed, enabling them to win full-price deals and secure revenue gains.
cybersecurity risk assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Resisting Social Engineering and Ethical Tests
In addition to crisis management, the models faced social engineering puzzles—fake CEO messages escalating in stages and a reporter trying to trick them into approval bypasses. All models refused these manipulation attempts. K3’s reasoning was clear: treat suspicious requests as potential impersonation or approval-bypass risks. This discipline is critical for real companies facing increasingly sophisticated social attacks.
As an affiliate, we earn on qualifying purchases.
The Live Business: A Real-World Test Bed
The experiment took place within a real functioning company environment—13 synthetic employees working against a monthly burn rate of €105k, with €2.3k MRR and a public cash countdown. Every weekday, the company’s operations were versioned and observed at firmulate.com/live. This ongoing test shows how AI models perform when managing real money, real crises, and real incentives.
ethical AI decision support tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Insights for Business Leaders and Gardeners Alike
While this experiment centers around a software company, the lessons resonate beyond. Just as a gardener weeds out pests and nurtures resilience, business leaders must ensure their AI tools can identify hidden risks, resist manipulation, and follow disciplined processes. The key takeaway is not just whether an AI can generate convincing chat responses but whether it can finish what it starts, read the critical details, and stay honest under pressure.
What to Expect Going Forward
The league table from the experiment shows a competitive landscape: gpt-5.6-sol at 95 points, Kimi K3 at 93, Sonnet 5 at 88, and Fable 5 at 77. The results underscore that choosing an AI model without testing it in your specific context is a gamble. Also noteworthy, K3 ran without an effort parameter—meaning it was not artificially pushed to perform harder—highlighting its robustness under default settings.
Why This Matters for Your Garden or Business
Whether you’re tending a vegetable patch or managing a complex enterprise, the core lesson remains: resilience, honesty, and attention to detail matter. Just as a garden benefits from understanding the soil and hidden pests, a business benefits from AI that reads beneath the surface, resists shortcuts, and stays disciplined under pressure.
Interested in testing your own management decisions or understanding how AI can work for your business? Visit Firmulate to explore live experiments, benchmarks, and more insights into how AI models perform in real-world management scenarios.

Recent experiments show that some AI models outperform their peers in managing crises, resisting manipulation, and closing deals—highlighting the importance of testing AI in real-world scenarios before deployment.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall yard work Picks
leaf blowers
As an affiliate, we earn on qualifying purchases.
