AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

In a world where artificial intelligence is becoming more embedded in daily business operations, the real question is not just about AI’s intelligence but about its integrity and discipline under pressure. A groundbreaking live experiment by Firmulate reveals that the top-performing AI models aren’t just good at processing data—they excel at decision-making in critical situations, often outperforming established expectations and even rivaling human judgment.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

The Live Wargame: Testing AI in a Real Business Crisis

Imagine a small software company facing its worst week—crises with customers, internal threats, and high-stakes negotiations—all happening simultaneously. In the Firmulate live experiment, four advanced AI models were tasked to run this company through the turmoil, making decisions just like human managers would. Every choice, from handling customer complaints to reading sensitive company files, was monitored, audited, and compared.

Measuring Trust and Performance

The results were striking. All four models identified every crisis and refused manipulation attempts—such as fake CEO messages designed to trick them into bypassing security protocols. Their responses kept the company’s integrity intact, demonstrating a level of discipline that is rare even among seasoned human managers.

The models’ ability to resist social engineering tactics stood out. When faced with staged impersonation attempts, all rejected the requests, with the Kimi K3 model explaining, “Treat the request as a suspected approval-bypass or possible impersonation.”

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Performance in Deal-Making and Data Reading

While all models performed well in crisis detection and resistance, only two managed to secure the company a lucrative deal worth €55,000, generating an additional €4,583 in monthly recurring revenue. Notably, the decisive factor was not superficial chat skills but the models’ ability to dig into company files—two references deep—uncover critical information that humans overlooked.

The Kimi K3 model, a newcomer from Moonshot, found the buried security data and successfully closed the deal at full price, outperforming peers that relied on surface-level analysis or slipped into process slips like writing attempts into a locked department instead of escalating.

Discipline Under Pressure

One telling example was how the models dealt with internal compliance lapses. The most thorough participant, Opus 4.8, with over 80 learned rules and deep analyses, surprisingly placed last in the final scores—only 73 out of 100. It left a close opportunity on the table and showed weaker discipline, demonstrating that thoroughness does not automatically translate into effective execution under stress.

The Fairness and Methodology

It’s essential to note that K3’s performance was evaluated with the default effort parameter, while others ran at a high effort setting (xhigh). This helps ensure that the comparison is fair and that K3’s strength is genuine, not a result of excessive tuning.

Implications for the Future of Business AI

This live experiment underscores a vital insight: success in AI-led decision-making isn’t just about generating responses that sound good in demos. It’s about consistency, honesty, and the ability to read and interpret complex data—skills that will determine whether AI can truly replace or augment human managers.

As firms consider integrating AI into their operations, the critical questions are: Will this AI finish what it starts? Will it read vital information hidden in files? Will it stay honest, even when tempted? These are the tests that matter in real-world decision-making, and the Firmulate experiment makes that clear.

Looking Beyond the Benchmarks

For more detailed results and to explore how these models performed across different scenarios, visit the Firmulate benchmarks page. This ongoing experiment is a live demonstration that choosing an AI model for your business isn’t just a leap of faith—it’s a strategic decision rooted in observable, measurable performance.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Powerful Tattoo Meanings for Mental Health

Fascinating tattoo symbols embody resilience and inner strength, offering powerful reminders in mental health battles – discover their profound meanings here.

The Triskelion: Triple Spiral Symbol in Celtic Art

Learn how the triskelion’s swirling design embodies Celtic beliefs in life, death, and rebirth, revealing a symbol of eternal transformation and resilience.

The Ouroboros: Snake Eating Its Tail in Art and Alchemy

Keen explorers of symbolism will find the ouroboros revealing profound insights into life’s eternal cycles and transformative power.

Why a Do-Nothing AI Baseline Scores 26 Points — and What It Means for Business Trustworthiness

A live AI management experiment reveals even a do-nothing baseline scores 26, emphasizing honesty and discipline as core to trustworthy AI performance in real-world business.