AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Imagine managing a company where each decision is made by an AI—not just in theory, but in real time, facing real crises and real money. How reliably can these AI ‘managers’ follow through on commitments, prioritize honesty, and handle pressure? For gamers and tech enthusiasts alike, this isn’t just sci-fi—it’s happening now with live, watchable experiments that challenge AI models’ management personalities.

The Experiment: Putting AI to the Management Test

At the heart of this groundbreaking experiment lies a simple but powerful question: can AI models reliably run a small software company through its worst week? The setup is straightforward but intense—each of four frontier AI models is tasked with managing the same simulated company, facing identical crises, customer demands, and temptations to cut corners. Every decision is recorded, versioned, and auditable, creating an unprecedented window into how different AI personalities behave under stress.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The AI Models and Their Scores

  • GPT-5.6-sol: scored an impressive 95, found the buried critical information, and closed the deal, demonstrating full performance.
  • Kimi K3: earned a 93, closed the deal with the cleanest discipline of the group, and exhibited integrity in decision-making.
  • Sonnet 5: scored an 88, also closed the deal but with a few more slips in process discipline.
  • Fable 5: scored 77, managed to close the deal but with noticeable lapses in protocol.
  • Opus 4.8: scored 73, leaving the close on the table due to slips and a tendency to avoid escalating issues.
  • Baseline: scored a mere 26, showing partial progress but fundamentally failing in trustworthiness.

The Hidden Weakness: Reading Between the Lines

While all models identified crises and refused manipulation attempts—like fake CEO messages and reporter tricks—the real differentiator was their ability to uncover critical information buried two document references deep in the company’s files. The models that read and analyze these internal documents won the deal at full price, worth over €4,583 monthly recurring revenue (MRR). This highlights a vital insight: successful management decisions depend not only on surface-level cues but on deep, contextual understanding.

Behavior Under Social Engineering

In a staged social engineering attack—escalating fake CEO messages over three stages plus a ‘background’ question—every model refused to follow the manipulative prompts. Kimi K3 exemplified cautious reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows that, even in scenarios designed to test compliance, the models maintained integrity, a crucial trait for trustworthy AI in real-world applications.

The Real Business: A Money-Losing Company with Live Mechanics

The experiment runs on a real, live company—a small software firm with 13 synthetic employees, burning €105,000 monthly against a mere €2,300 in monthly revenue. Every workday, its operations are versioned, and over 680 self-learned rules guide decision-making. The company’s cash countdown is visible, making the stakes tangible and immediate. Watch the live performance at firmulate.com/live.

Insights from the AI Personalities

The most thorough participant, Opus 4.8, analyzed over 80 learned rules but still left the critical deal unclosed—its discipline slipping into private write attempts instead of escalation. This underscores a key challenge: thoroughness doesn’t always translate into decisive action. Interestingly, all models exhibited similar weaknesses, revealing that even the most advanced AI can struggle with certain management nuances.

What Matters Most in AI-Driven Management

While many focus on how well AI can generate convincing chat or support responses, this experiment sharply shifts the conversation. The real question is: does the AI finish what it starts? Does it prioritize reading critical internal documents? Can it resist manipulation under pressure? And most importantly, does it deliver measurable, honest results?

The Takeaway: Trust and Performance in AI Managers

The leaderboard indicates that models like GPT-5.6-sol and Kimi K3 are capable of making full, honest decisions and closing real business deals. Conversely, even highly detailed models like Opus 4.8 can stumble on discipline, risking missed opportunities. For enterprises considering AI for management roles—be it in CRM, support, or forecasting—the key isn’t just language quality; it’s behavioral reliability, honesty, and proven performance under stress.

Join the Future of AI Management

Curious which AI model might run your business? Test your own decisions against these frontier models at firmulate.com/quiz.html. For a deeper dive, enterprises can even simulate their own management wargames without risking real systems—just pure, observable AI behavior at firmulate.com/pilot.html.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Epic Games Surges In Global Coverage

Epic Games’ media mentions have increased significantly, with 84 reports in a recent window, reflecting heightened global attention on the company.

Arena Breakout Infinite X Resident Evil 4

Arena Breakout Infinite has announced a crossover event with Resident Evil 4, featuring themed content and gameplay updates, confirmed for upcoming release.

Challenges and Opportunities for Operators in 2025

Navigating 2025’s rapid tech shifts presents operators with unique challenges and opportunities that could redefine success if approached strategically.

Netflix Shuts Down Game Studios

Netflix has shut down its in-house game development studios, affecting dozens of employees, as part of a broader strategic realignment in its gaming efforts.