
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
Gaming the System or Playing It Straight? AI’s True Test in Business Decision-Making
Just as gamers test the limits of strategy and honesty in virtual worlds, AI models are now put through their paces in real-world business simulations. The question isn’t just about their intelligence or speed—it’s about integrity, focus, and the ability to finish what they start. In an unprecedented live experiment, four AI models faced off in managing a small software company’s worst week—complete with crises, temptations, and high-stakes decisions.
The Setup: Simulating a Business Crisis
In this experiment, each AI model was tasked with running the same small software company through its most challenging week. They faced identical customer issues, internal crises, and were tempted with manipulative tactics designed to test their integrity. Every decision made was recorded, versioned, and auditable, ensuring transparency and accountability.
The models included the latest frontier AIs—gpt-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8—each with unique strengths and approaches. Notably, Opus 4.8 stood out for its thoroughness, having learned over 80 rules to guide decision-making, and delivering deep analyses. Yet, it finished last in the final results.
Key Findings: Honesty and Focus Trump Diligence
All four models successfully identified every crisis and refused every attempt at manipulation, including sophisticated social engineering schemes like fake CEO messages and reporter tricks. This demonstrates that even the most thorough AI, such as Opus 4.8, can be swayed by sheer volume of learned rules if discipline falters.
In a critical turn, only two models—Kimi K3 and gpt-5.6-sol—secured the full €55,000 deal, having made the same diagnosis and pitch as their competitors. The difference? They read deeper into the company’s own files, uncovering a crucial fact buried two document references deep in the company’s data. This insight was decisive and secured a full subscription worth +€4,583 in monthly recurring revenue (MRR).
The Hidden Weakness: Superficial Analysis vs. Deep Reading
Despite their diligence and ability to recognize crises, the most thorough participant—Opus 4.8—missed the deal because, during the close, discipline slipped. Instead of escalating a critical issue, it wrote attempts into a locked department, missing an opportunity to fully close the deal. This pattern was observed across all models in varying degrees: volume of learned rules does not guarantee impact without disciplined prioritization.
The Real-World Implications
In a live setting with 13 synthetic employees and real money mechanics—burning €105k monthly against a €2.3k MRR—these insights matter. The models are tested in real-time, managing actual cash flow and business processes, with every decision versioned and transparent.
The takeaway is clear: for AI to be truly effective in critical business roles, it must do more than identify problems and refuse manipulative tactics. It must prioritize effectively, read deeply into relevant data, and stay disciplined under pressure. Diligence without focus can still lead to missed opportunities, even for the most thorough AI models.

Key Takeaway
In high-stakes business decisions, AI’s impact hinges less on volume of rules or learned knowledge and more on disciplined prioritization and deep data reading. The most thorough AI can still lose if it slips on focus—emphasizing that quality of work beats quantity in AI-driven enterprise management.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
business AI decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.