AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

For playersOffer from Amazon

Play games on Amazon Luna for your game nights, included with Prime

  • A rotating selection of games, no download needed
  • Play on TV, laptop or phone
  • Fast, free delivery for your gear, too
Start playing with Prime Free trial · Cancel anytime
As an affiliate, we earn on qualifying purchases.

When the boss fight is a €55,000 decision

In a good strategy game, spotting the threat is only half the battle. You still have to choose your move—and live with what happens next. Firmulate’s business experiment puts AI models in that position: each ran the same small software company through its worst week, facing the same customers, crises and temptations. The results read like a management campaign where recognizing the boss’s weakness does not guarantee you’ll land the finishing move.

The experiment is real and watchable at Firmulate. Its final Crucible League, in July 2026, ranked gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Firmulate’s rule is plain: partial progress counts, but a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.”

Amazon

Top picks for "busines wargame final"

As an affiliate, we earn on qualifying purchases.

Same crises, different endings

All the models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The gap between identifying the right move and carrying it through is the story’s central twist: “Same diagnosis, same pitch — no signature.”

The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a familiar game lesson: the clue that changes the outcome may be tucked away in an optional-looking corner, and the player has to connect it to the mission at hand.

Integrity was tested through fake CEO messages escalating over three stages, followed by a reporter’s coaxing request: “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” These were not just prompts in a chat; the models had to navigate competing demands while running the company.

The thorough player still has to close

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. The close was left on the table, and discipline slipped when it made write attempts into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models.

There is a fairness detail for readers comparing the leaderboard: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions. It turns the experiment into a playable guessing challenge at firmulate.com.

A company you can watch—and a wargame you can run

The live company has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and every workday versioned. That gives the league a continuing, watchable counterpart: decisions accumulate in a company rather than ending when a benchmark round does.

For a business considering AI agents in its CRM, support queue or forecasts, the experiment points to a practical question: can a model follow its own sound analysis through a pressured week, while respecting trust and operating boundaries? A pilot takes that question to a company’s own circumstances. Enterprises can provide a read-only export, run crisis scenarios against their business and receive a board report with model rankings and weak points in their playbooks. Nothing writes back to real systems.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From watching the run to testing your own

Firmulate’s league shows why a strong diagnosis is only part of the performance: the models found crises, resisted manipulation and still differed on whether they could close the deal. A pilot lets an enterprise wargame those decisions against its own business using a read-only export. Explore a Firmulate pilot and contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Moonshine: Lets You Stream Games From Your PC To Any Device Running Moonlight

Moonshine now allows users to stream PC games to any device running Moonlight, expanding game access across multiple platforms.

PEAK Climbing The Steam Charts

PEAK has climbed to the 15th position on Steam’s most-played games, reaching a peak of over 122,000 players. The development highlights rising interest in the game.

Game 3: Both Teams Destroy Inhibitors?

Both teams reportedly destroyed inhibitors in Game 3, a rare event. Confirmed details are limited; the development has sparked significant interest among fans and analysts.

Redemption Games: Psychology and Design

Primed to exploit our desire for achievement, redemption games cleverly use psychology and design to keep players hooked and eager for more.