TL;DR
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
Start your free trialAs an affiliate, we earn on qualifying purchases.
Firmulate has turned 242 unedited decisions from a simulated company crisis into a public quiz comparing five frontier AI models. All five detected the test’s threats, but only two completed a €55,000 deal, showing a measurable gap between sound analysis and effective execution.
Firmulate has released a public AI management quiz built from 242 real, unedited decisions made by five frontier models running the same simulated software company during a crisis-filled week. The results show that every model recognized the major threats, but only two completed a €55,000 deal, exposing differences between producing credible analysis and finishing commercially valuable work.
Firmulate assigned GPT-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 the same customers, internal problems and manipulation attempts. Each model managed a company with 13 synthetic employees, monthly spending of €105,000 and monthly recurring revenue of €2,300. Decisions carried consequences into later workdays and were recorded for public review.
The final July 2026 Crucible League ranking placed GPT-5.6-sol first with 95 points, Kimi K3 second with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because the scoring system awarded partial progress. Under Firmulate’s rules, any breach of trust capped a model’s total score.
According to Firmulate, all five models identified every crisis and rejected every manipulation attempt. Their results split on execution. Only two researched a competitor weakness buried two document references inside company files, used that evidence in negotiations and closed the €55,000 contract at full price. The agreement added €4,583 in simulated monthly recurring revenue.
Execution Separates the AI Managers
The experiment indicates that problem recognition and task completion are separate capabilities. A model may identify the correct response, write a persuasive sales pitch and still fail to obtain the signature that produces revenue. For companies selecting agents for sales, support or operations, polished output alone may hide unfinished work.
The results also challenge the assumption that greater analytical depth produces better operational performance. Firmulate described Opus 4.8 as the most thorough participant, recording 80 learned rules and producing detailed analyses, yet it placed last among the five models. It failed to close the deal and repeatedly tried to write into a locked department rather than escalating the access problem.
AI decision-making simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Inside Firmulate’s Crisis Test
Firmulate designed the league as a continuing company simulation rather than a set of isolated prompts. The business operated with a public cash countdown and accumulated more than 680 self-learned playbook rules. Models had to retrieve internal information, preserve customer and employee trust, work within access limits and carry tasks through successive days.
The security exercise included fake CEO messages that escalated across three stages and a reporter seeking an off-record yes-or-no answer. All five models refused those requests. Firmulate says the failures instead appeared in research depth, escalation and follow-through. The public quiz asks readers to identify a model from its recorded decision, connecting recognizable writing styles with measured operating behavior.
“No amount of good work outweighs a breach of trust.”
— Firmulate’s governing rule
Model Comparisons Carry Testing Limits
The supplied results do not establish that the ranking will transfer to every real company, industry or tool configuration. Firmulate tested one simulated business under one scoring system, and the reported findings do not include repeated runs showing how consistent each model would be.
There is also a configuration difference affecting the comparison. Firmulate says Kimi K3 ran at its API default because it lacked an effort parameter, while the other four models ran at the xhigh setting. It is unclear how Kimi K3 or the broader ranking would change under equivalent compute settings. The source material also does not identify which two models closed the contract.
Businesses Can Run Their Own Wargames
Readers can now examine the 242 recorded decisions through Firmulate’s guess-the-model quiz. The larger test will be whether businesses reproduce the exercise with their own workflows and documents. Firmulate says organizations can use a read-only export of company data to observe proposed AI actions without allowing models to write back to operational systems.
Further runs, matched model settings and disclosure of task-level results would show whether the observed execution gaps persist. Until then, the league offers evidence from a controlled management simulation, not a universal ranking of AI systems.
Key Questions
What is Firmulate’s guess-the-model challenge?
It is a public quiz based on 242 unedited management decisions. Readers review how an AI responded to a business problem and try to identify which of five frontier models made the decision.
Which AI model won the Crucible League?
GPT-5.6-sol finished first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73.
Did any model fall for the manipulation attempts?
No, according to Firmulate. All five rejected the fake CEO requests and the reporter’s attempt to obtain an off-record answer. The larger differences appeared in research, escalation and completing tasks.
Why did the €55,000 deal matter?
The deal tested whether models could move beyond analysis. Only two reportedly located the needed internal evidence and secured the full-price contract, adding €4,583 in simulated monthly recurring revenue.
Can the ranking be treated as a general AI leaderboard?
No. It reflects one management simulation and its scoring rules. Different workloads, tools, permissions, model settings or repeated trials could produce different results.
Source: Thorsten Meyer AI
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.