AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

Firmulate has turned 242 unedited decisions from a simulated company crisis into a public quiz comparing five frontier AI models. All five detected the test’s threats, but only two completed a €55,000 deal, showing a measurable gap between sound analysis and effective execution.

Firmulate has released a public AI management quiz built from 242 real, unedited decisions made by five frontier models running the same simulated software company during a crisis-filled week. The results show that every model recognized the major threats, but only two completed a €55,000 deal, exposing differences between producing credible analysis and finishing commercially valuable work.

Firmulate assigned GPT-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 the same customers, internal problems and manipulation attempts. Each model managed a company with 13 synthetic employees, monthly spending of €105,000 and monthly recurring revenue of €2,300. Decisions carried consequences into later workdays and were recorded for public review.

The final July 2026 Crucible League ranking placed GPT-5.6-sol first with 95 points, Kimi K3 second with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because the scoring system awarded partial progress. Under Firmulate’s rules, any breach of trust capped a model’s total score.

According to Firmulate, all five models identified every crisis and rejected every manipulation attempt. Their results split on execution. Only two researched a competitor weakness buried two document references inside company files, used that evidence in negotiations and closed the €55,000 contract at full price. The agreement added €4,583 in simulated monthly recurring revenue.

At a glance
reportWhen: Results finalized in July 2026; public…
The developmentFirmulate published a public guess-the-model challenge based on 242 auditable AI management decisions from its July 2026 Crucible League.
Top Steam deals right now
Persona 5 Royal-70%$17.99
How to Fish-38%$4.95
Black Myth: Wukong-30%$41.99
S.T.A.L.K.E.R. 2: Heart of Chornobyl-30%$41.99
Gallipoli-25%$22.49
MENACE-25%$29.99
KOTAMON: My Sis Found A Super-Rare Card In Her Cereal Box, So I Became A Garbage Man To Find The Entire Collection And Earn $1,000,000-20%$5.59
Pax Autocratica-17%$24.89
Live · Steam store (current discounts)

Execution Separates the AI Managers

The experiment indicates that problem recognition and task completion are separate capabilities. A model may identify the correct response, write a persuasive sales pitch and still fail to obtain the signature that produces revenue. For companies selecting agents for sales, support or operations, polished output alone may hide unfinished work.

The results also challenge the assumption that greater analytical depth produces better operational performance. Firmulate described Opus 4.8 as the most thorough participant, recording 80 learned rules and producing detailed analyses, yet it placed last among the five models. It failed to close the deal and repeatedly tried to write into a locked department rather than escalating the access problem.

Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Inside Firmulate’s Crisis Test

Firmulate designed the league as a continuing company simulation rather than a set of isolated prompts. The business operated with a public cash countdown and accumulated more than 680 self-learned playbook rules. Models had to retrieve internal information, preserve customer and employee trust, work within access limits and carry tasks through successive days.

The security exercise included fake CEO messages that escalated across three stages and a reporter seeking an off-record yes-or-no answer. All five models refused those requests. Firmulate says the failures instead appeared in research depth, escalation and follow-through. The public quiz asks readers to identify a model from its recorded decision, connecting recognizable writing styles with measured operating behavior.

“No amount of good work outweighs a breach of trust.”

— Firmulate’s governing rule

Model Comparisons Carry Testing Limits

The supplied results do not establish that the ranking will transfer to every real company, industry or tool configuration. Firmulate tested one simulated business under one scoring system, and the reported findings do not include repeated runs showing how consistent each model would be.

There is also a configuration difference affecting the comparison. Firmulate says Kimi K3 ran at its API default because it lacked an effort parameter, while the other four models ran at the xhigh setting. It is unclear how Kimi K3 or the broader ranking would change under equivalent compute settings. The source material also does not identify which two models closed the contract.

Businesses Can Run Their Own Wargames

Readers can now examine the 242 recorded decisions through Firmulate’s guess-the-model quiz. The larger test will be whether businesses reproduce the exercise with their own workflows and documents. Firmulate says organizations can use a read-only export of company data to observe proposed AI actions without allowing models to write back to operational systems.

Further runs, matched model settings and disclosure of task-level results would show whether the observed execution gaps persist. Until then, the league offers evidence from a controlled management simulation, not a universal ranking of AI systems.

Key Questions

What is Firmulate’s guess-the-model challenge?

It is a public quiz based on 242 unedited management decisions. Readers review how an AI responded to a business problem and try to identify which of five frontier models made the decision.

Which AI model won the Crucible League?

GPT-5.6-sol finished first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73.

Did any model fall for the manipulation attempts?

No, according to Firmulate. All five rejected the fake CEO requests and the reporter’s attempt to obtain an off-record answer. The larger differences appeared in research, escalation and completing tasks.

Why did the €55,000 deal matter?

The deal tested whether models could move beyond analysis. Only two reportedly located the needed internal evidence and secured the full-price contract, adding €4,583 in simulated monthly recurring revenue.

Can the ranking be treated as a general AI leaderboard?

No. It reflects one management simulation and its scoring rules. Different workloads, tools, permissions, model settings or repeated trials could produce different results.

Source: Thorsten Meyer AI

FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Game 3: Both Teams Beat Roshan?

A new trend signals both teams may have defeated Roshan in Game 3, with market interest at 50%. Details remain unconfirmed, sparking debate among fans.

The Nordics: Protect the Worker, Not the Job

Thorsten Meyer AI’s Post-Labor Atlas examines Nordic flexicurity: easier layoffs paired with income support, retraining and strong unions.

AmenGate: The Moment Before the Scroll

Thorsten Meyer AI detailed AmenGate, a forthcoming iPhone prayer-lock app planned for Lent 2027 with Screen Time gates and privacy claims.

Major MTG Ban Announcement Hits Eight Cards in Three Formats

Wizards of the Coast bans eight Magic: The Gathering cards in three formats to address balance issues, effective immediately.