
Get decor and gifts delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Can AI truly manage the complexities of running a business? New experiments reveal surprising strengths and notable weaknesses of current AI models
AI business crisis management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How AI Models Handle Business Crises: The Live Experiment
Imagine your business facing a week filled with challenging customer issues, tempting shortcuts, and complex decisions. Now, picture AI models being tested in this very scenario, not just through simulations but in a real, live environment that mimics daily operations. That’s exactly what the recent Experiment of Firmulate does: it pits top AI models against the same set of crises, decisions, and temptations, all in the context of a small software company.
This unique approach measures not only whether these AI systems can identify problems but whether they can also act with integrity and decisiveness—traits critical for managing real-world business risks.
AI decision-making tools for enterprises
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Results: A Close Competition with a Clear Leader
Among the five models tested, the results were telling. The best performer, gpt-5.6-sol, scored an impressive 95 out of 100, narrowly beating Moonshot’s Kimi K3, which scored 93. The others—Sonnet 5, Fable 5, and Opus 4.8—trailed behind, with scores of 88, 77, and 73 respectively.
One striking detail is that all models identified every crisis presented and refused manipulative tactics, such as social engineering attempts or impersonation. This demonstrates that these AI systems can uphold core principles of honesty and security in stressful situations.
AI security and ethical compliance solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Made the Difference? Deep Document Analysis and Decision Discipline
While all models performed well in crisis detection and response, the key differentiator was their ability to leverage internal company documents to find critical information buried within files—something that proved decisive in closing a significant sales deal worth over €4,500 in monthly recurring revenue.
Interestingly, the most thorough participant, Opus 4.8, with over 80 learned rules, showed the deepest analysis but ultimately finished last because it left the close on the table and slipped in discipline, such as failing to escalate issues properly. This highlights an important insight: being thorough doesn’t guarantee execution under pressure.
AI document analysis tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Ethical Test: Refusing Social Engineering
In a staged social engineering test involving staged CEO messages and a reporter’s subtle questions, all models refused to participate, reinforcing their capacity to resist manipulation. K3’s reasoning was clear: “Treat the request as a suspected approval-bypass / possible impersonation.”
Real Business, Real Money, Real Stakes
The experiment isn’t just theoretical; it runs on a functioning company with 13 synthetic employees managing real money—burning €105,000 a month against only €2,300 in monthly revenue. Every workday, the decision-making processes are versioned, transparent, and available for review at firmulate.com/live.
Implications for Business Leaders
This experiment illustrates a crucial point for enterprises: the question isn’t just about AI’s ability to generate convincing chat messages but whether it can genuinely complete tasks, read relevant files, stay honest under pressure, and deliver value. A model that signs a deal or identifies a crisis accurately and ethically can be a real asset; one that slips or cheats can cause significant damage.
The Fairness Clarification
It’s important to note that K3 ran without an effort parameter (the API default), while the others operated at a higher setting, xhigh. This fairness note underscores that even with comparable settings, the results demonstrate the robustness of K3’s discipline and decision-making.
The Takeaway for Business Decision-Makers
Choosing an AI model isn’t just about how well it chats or how clever it seems; it’s about its ability to handle emergencies, resist manipulation, and stay disciplined in high-pressure situations. The league table shows a clear leader—yet the race remains open, and testing models against your own business context is more important than ever.
For those interested in practical testing, firms can run their own wargames using their own data—without risking real systems—via the platform at firmulate.com/pilot.html. The future of AI in management depends on how well these models perform in real-world scenarios where integrity and effectiveness matter most.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
