AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine a garden where the strongest plants survive not by just growing tall but by resisting pests, drought, and disease. Now, picture managing a business with the same resilience. Recent experiments with artificial intelligence (AI) show that some models can outperform human decision-making in managing a company’s toughest week — refusing temptation, uncovering hidden risks, and sealing profitable deals. It’s a surprising story for anyone interested in sustainable growth, whether you’re tending a garden or growing a business.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get garden gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Experiment: Simulating a Business Crisis

Four advanced AI models faced the same challenge: run a small software company through its worst week. This wasn’t a simple test of chat skills; it was a rigorous simulation involving real crises—customer churn, security breaches, and manipulative social engineering attempts. Every decision was carefully tracked and auditable, mimicking real management decision-making. The goal was to see which model could identify risks, resist temptation, and close profitable deals under pressure.

Results that Defy Expectations

All four models successfully identified every crisis and refused manipulation attempts. This demonstrates a remarkable capacity for honesty and discernment—key traits for managing real business risks. However, only two of the four models managed to close a €55,000 deal that their own analysis had earned them. The difference? One model found a buried document reference that revealed a crucial security flaw—allowing it to close the deal at full price, adding +€4,583 MRR.

The Surprise Winner: The Newcomer

The top scorer, gpt-5.6-sol, led with a score of 95 out of 100. Close behind was Kimi K3 from Moonshot, scoring 93. Despite being a newcomer, K3 demonstrated the cleanest discipline—resisting all baits and only deviating once in a disciplined manner. The other models, Sonnet 5 and Fable 5, scored 88 and 77 respectively, with Opus 4.8 trailing at 73. They all managed the crises but showed weaknesses in process discipline and risk detection.

Amazon

AI business decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Edge: Reading Beyond the Surface

The crucial edge for K3 wasn’t just in crisis diagnosis but in reading deeper into company files. The experiment revealed that the decisive advantage lay two document references deep in internal files, not in external customer interactions. Models that read and analyze these internal documents were able to uncover risks and seize opportunities that others missed, enabling them to win full-price deals and secure revenue gains.

Amazon

cybersecurity risk assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Resisting Social Engineering and Ethical Tests

In addition to crisis management, the models faced social engineering puzzles—fake CEO messages escalating in stages and a reporter trying to trick them into approval bypasses. All models refused these manipulation attempts. K3’s reasoning was clear: treat suspicious requests as potential impersonation or approval-bypass risks. This discipline is critical for real companies facing increasingly sophisticated social attacks.

Amazon

business crisis management AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Live Business: A Real-World Test Bed

The experiment took place within a real functioning company environment—13 synthetic employees working against a monthly burn rate of €105k, with €2.3k MRR and a public cash countdown. Every weekday, the company’s operations were versioned and observed at firmulate.com/live. This ongoing test shows how AI models perform when managing real money, real crises, and real incentives.

Amazon

ethical AI decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Insights for Business Leaders and Gardeners Alike

While this experiment centers around a software company, the lessons resonate beyond. Just as a gardener weeds out pests and nurtures resilience, business leaders must ensure their AI tools can identify hidden risks, resist manipulation, and follow disciplined processes. The key takeaway is not just whether an AI can generate convincing chat responses but whether it can finish what it starts, read the critical details, and stay honest under pressure.

What to Expect Going Forward

The league table from the experiment shows a competitive landscape: gpt-5.6-sol at 95 points, Kimi K3 at 93, Sonnet 5 at 88, and Fable 5 at 77. The results underscore that choosing an AI model without testing it in your specific context is a gamble. Also noteworthy, K3 ran without an effort parameter—meaning it was not artificially pushed to perform harder—highlighting its robustness under default settings.

Why This Matters for Your Garden or Business

Whether you’re tending a vegetable patch or managing a complex enterprise, the core lesson remains: resilience, honesty, and attention to detail matter. Just as a garden benefits from understanding the soil and hidden pests, a business benefits from AI that reads beneath the surface, resists shortcuts, and stays disciplined under pressure.

Interested in testing your own management decisions or understanding how AI can work for your business? Visit Firmulate to explore live experiments, benchmarks, and more insights into how AI models perform in real-world management scenarios.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Recent experiments show that some AI models outperform their peers in managing crises, resisting manipulation, and closing deals—highlighting the importance of testing AI in real-world scenarios before deployment.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Beachcomber’s Basket: Foraging Coastal Succulents for Crisp Coastal Salads

Nestled along the shoreline, discover how foraging coastal succulents can elevate your salads with salty, crunchy flavors—read on to learn more.

15 Wild Edibles to Forage for Food, According to Expert Foragers

Discover the top foraging guides and resources for edible wild foods in 2026. Find the best picks for beginners, experts, and safety-focused foragers.

The Moon Is Running the Beach: How Lunar Cycles Change Coastal Harvests

B**e prepared to uncover how lunar cycles influence coastal harvests and discover secrets that could transform your approach to harvesting.

14 Best Tips on How to Forage for Beginners

Discover the top foraging guides for beginners in 2026. Find the best overall, budget-friendly, and beginner-friendly options to start your wild food adventure.