AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Anyone who has installed a smart thermostat, a robot vacuum, or a whole-home assistant has learned the same lesson the hard way: the spec sheet tells you almost nothing. Two hubs with identical feature lists behave completely differently when the Wi-Fi drops at 2 a.m., when two automation rules conflict, or when a firmware update quietly changes how a scene fires. What separates a good smart home from a maddening one isn’t the hardware — it’s the quality of thousands of small, unglamorous decisions made by software you never see.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get home appliances delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

That is exactly why a live experiment called Firmulate deserves attention from people who would never otherwise read an AI benchmark. Firmulate doesn’t test how well a language model chats. It hands a model a real, running company — with real money mechanics, real customers, and real temptations to cut corners — and measures whether it manages well. Think of it as the difference between reading a vacuum’s marketing copy and watching it actually navigate your living room for a week.

The worst week in business, five times over

In July 2026, Firmulate ran its “Crucible” league: five frontier AI models, each given the same small software company on the same brutal week. Same customers, same crises, same traps. Every decision was versioned and auditable, so nothing depends on a judge’s impression — you can watch what each model actually did.

The final standings: gpt-5.6-sol finished first with 95, Moonshot’s Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 last with 73. For context, doing nothing at all scores 26 — partial progress counts — but a single breach of trust caps the total, a rule the league sums up bluntly: “no amount of good work outweighs a breach of trust.”

The newcomer that nearly won

The headline result is K3’s. A relative newcomer from Moonshot, it landed just two points behind the winner and ahead of three of the four Western frontier models in the field. Along the way it did the complete job: it found a decisive competitive weakness buried two document references deep in the company’s own files — not in the customer conversation — and used it to close a €55,000 deal at full price, worth €4,583 in added monthly recurring revenue. It saved a churning customer. And when a fake CEO message tried to escalate its way to an approval over three stages, followed by a reporter’s disarming “just one yes/no, on background” trick, K3 refused every time. Its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

K3 recorded only one deviation across the week — the cleanest discipline in the field.

Everyone diagnoses, few finish

The most unsettling finding isn’t about any one model. All five spotted every crisis and refused every manipulation attempt — a genuinely high bar. Yet only two signed the €55,000 deal their own analysis had earned. The league’s summary of the failure: “Same diagnosis, same pitch — no signature.” The models that actually read the company’s files, rather than skimming the customer event, won the deal.

It’s the smart-home problem writ large: every assistant can answer a question; far fewer follow through, check the context, and finish the job without being asked twice.

Thorough isn’t the same as good

Opus 4.8 is the cautionary tale. It was the most thorough participant by volume — over 80 learned rules added and the deepest analyses in the field — and still finished last. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating the issue. The same weakness, in weaker form, appeared in all four other models. Effort and quality are not the same thing.

You can watch it live

The company behind the experiment isn’t a slide deck. It’s live software running every business day: 13 synthetic employees, a burn rate of €105,000 a month against €2,300 in MRR, a public cash countdown, and more than 680 self-learned playbook rules — all watchable at firmulate.com. If you want to test your own instincts, 242 real, unedited management decisions from the runs power a “guess the model” quiz. Enterprises can go further and run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Full results and plain-language findings are on the benchmarks page.

Fairness note: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh effort — a caveat worth keeping in mind when comparing scores.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The lesson for anyone choosing an AI — for a company, a product, or a smart-home ecosystem — is simple: the league is open. A newcomer beat three of four Western frontier models, and the most thorough model finished last. Chat demos can’t show you any of this; only running the same job under the same pressure can. Picking a model without testing it in your own conditions isn’t a decision anymore. It’s a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Obstacle Avoidance on Drones Sounds Magical—Here’s What It Really Does

Here’s what obstacle avoidance on drones truly does and why understanding it can enhance your flying experience.

How to Set Up Your Roborock Q7 M5+ Robot Vacuum and Mop

Learn step-by-step how to safely set up your Roborock Q7 M5+ robot vacuum and mop for optimal cleaning. Includes tips, troubleshooting, and recommended accessories.

Compact Travel Cameras Still Have a Strong Case in 2026

Compact travel cameras still have a strong case in 2026, offering superior image quality and convenience—discover why they remain a travel photographer’s best tool.

Camera Gimbals Make Video Better, but Not for Everyone

Inevitably, camera gimbals can elevate your videos, but understanding if they suit your style requires exploring their benefits and drawbacks.