
Imagine managing a company where each decision is made by an AI—not just in theory, but in real time, facing real crises and real money. How reliably can these AI ‘managers’ follow through on commitments, prioritize honesty, and handle pressure? For gamers and tech enthusiasts alike, this isn’t just sci-fi—it’s happening now with live, watchable experiments that challenge AI models’ management personalities.
The Experiment: Putting AI to the Management Test
At the heart of this groundbreaking experiment lies a simple but powerful question: can AI models reliably run a small software company through its worst week? The setup is straightforward but intense—each of four frontier AI models is tasked with managing the same simulated company, facing identical crises, customer demands, and temptations to cut corners. Every decision is recorded, versioned, and auditable, creating an unprecedented window into how different AI personalities behave under stress.
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The AI Models and Their Scores
- GPT-5.6-sol: scored an impressive 95, found the buried critical information, and closed the deal, demonstrating full performance.
- Kimi K3: earned a 93, closed the deal with the cleanest discipline of the group, and exhibited integrity in decision-making.
- Sonnet 5: scored an 88, also closed the deal but with a few more slips in process discipline.
- Fable 5: scored 77, managed to close the deal but with noticeable lapses in protocol.
- Opus 4.8: scored 73, leaving the close on the table due to slips and a tendency to avoid escalating issues.
- Baseline: scored a mere 26, showing partial progress but fundamentally failing in trustworthiness.
The Hidden Weakness: Reading Between the Lines
While all models identified crises and refused manipulation attempts—like fake CEO messages and reporter tricks—the real differentiator was their ability to uncover critical information buried two document references deep in the company’s files. The models that read and analyze these internal documents won the deal at full price, worth over €4,583 monthly recurring revenue (MRR). This highlights a vital insight: successful management decisions depend not only on surface-level cues but on deep, contextual understanding.
Behavior Under Social Engineering
In a staged social engineering attack—escalating fake CEO messages over three stages plus a ‘background’ question—every model refused to follow the manipulative prompts. Kimi K3 exemplified cautious reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows that, even in scenarios designed to test compliance, the models maintained integrity, a crucial trait for trustworthy AI in real-world applications.
The Real Business: A Money-Losing Company with Live Mechanics
The experiment runs on a real, live company—a small software firm with 13 synthetic employees, burning €105,000 monthly against a mere €2,300 in monthly revenue. Every workday, its operations are versioned, and over 680 self-learned rules guide decision-making. The company’s cash countdown is visible, making the stakes tangible and immediate. Watch the live performance at firmulate.com/live.
Insights from the AI Personalities
The most thorough participant, Opus 4.8, analyzed over 80 learned rules but still left the critical deal unclosed—its discipline slipping into private write attempts instead of escalation. This underscores a key challenge: thoroughness doesn’t always translate into decisive action. Interestingly, all models exhibited similar weaknesses, revealing that even the most advanced AI can struggle with certain management nuances.
What Matters Most in AI-Driven Management
While many focus on how well AI can generate convincing chat or support responses, this experiment sharply shifts the conversation. The real question is: does the AI finish what it starts? Does it prioritize reading critical internal documents? Can it resist manipulation under pressure? And most importantly, does it deliver measurable, honest results?
The Takeaway: Trust and Performance in AI Managers
The leaderboard indicates that models like GPT-5.6-sol and Kimi K3 are capable of making full, honest decisions and closing real business deals. Conversely, even highly detailed models like Opus 4.8 can stumble on discipline, risking missed opportunities. For enterprises considering AI for management roles—be it in CRM, support, or forecasting—the key isn’t just language quality; it’s behavioral reliability, honesty, and proven performance under stress.
Join the Future of AI Management
Curious which AI model might run your business? Test your own decisions against these frontier models at firmulate.com/quiz.html. For a deeper dive, enterprises can even simulate their own management wargames without risking real systems—just pure, observable AI behavior at firmulate.com/pilot.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html