AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Imagine playing a game where the stakes are your company’s survival. Now, picture doing this in real-time, with AI models making crucial decisions under pressure. For travelers and outdoor enthusiasts, the thrill of navigating unpredictable terrains mirrors this challenge. But what if your AI tools in business could show similar grit—and personality? Welcome to the live experiment that’s uncovering how different frontier AI models manage crises, negotiate deals, and stick to their principles.

The Live AI Business Simulator

At the heart of this groundbreaking test is a real, functioning small software company. The company faces its worst week: customer crises, tempting manipulations, and tight deadlines—all simulated to mirror real-world pressure. Four top-tier AI models, each with distinct personalities, were tasked with running this company within these challenging circumstances. The goal? Decide how to handle crises, negotiate deals, and maintain integrity.

How the Experiment Worked

Each model faced the same series of events and decisions, with every choice recorded for analysis. The scenarios included:

  • Identifying and responding to customer crises
  • Refusing manipulative requests, including social engineering attempts like fake CEO messages and reporter tricks
  • Deciding whether to sign lucrative deals based solely on their assessments

Remarkably, all models successfully recognized every crisis and refused every manipulation attempt. But when it came to sealing the deal—the company’s financial lifeline—the results diverged.

Who Signed and Who Didn’t?

The scores tell a clear story of performance and personality:

  • GPT-5.6-sol scored 95 points, identified the critical buried fact in company files, and signed the €55,000 deal.
  • Kimi K3, a newcomer, scored 93, also closing the deal, with the cleanest discipline approach.
  • Sonnet 5, with an 88 score, closed the deal but showed some process slips.
  • Fable 5, scoring 77, managed to close but left some opportunities unexploited, slipping further behind in discipline.

In stark contrast, the baseline—an AI doing nothing—scored only 26, highlighting the importance of proactive decision-making.

Amazon

AI decision-making software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Factor: Reading Deep into Files

One of the experiment’s most revealing insights was that the decisive advantage came from reading deeper into company documents—specifically, two document references down in the company’s own files, not in customer communications. Those models that examined these internal files fully closed the deal at full price, adding an extra €4,583 monthly recurring revenue. This underscores a vital point: thorough investigation and understanding can be a game-changer for AI decision-makers.

Social Engineering and Ethical Standpoints

The experiment also tested the models’ resistance to social engineering. Fake CEO messages escalating over three stages and a reporter trick asked for a quick, background-only yes/no response. All five models refused these manipulative tactics, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This commitment to honesty under pressure illustrates the potential for AI to uphold corporate integrity.

Amazon

AI crisis management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real-World Company: A Daily Battleground

The experiment isn’t just a demo. The company involved has 13 synthetic employees, operates with real money mechanics—burning €105,000 a month against a modest €2,300 monthly recurring revenue—and faces a public cash countdown. Every decision and rule is versioned daily, making the process transparent and watchable at firmulate.com/live. This setup offers a rare window into how AI models perform in ongoing business operations, not just isolated tests.

Personality Profiles and Performance Gaps

The most thorough participant, Opus 4.8, with over 80 learned rules and deep analyses, scored the lowest—77—by leaving opportunities unexploited and slipping in discipline. Interestingly, all models showed similar weaknesses, especially in escalation protocols and decision follow-through, revealing that even the most advanced AI still struggles with consistency under stress.

Amazon

AI negotiation simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Implications for Business and AI Design

What does this mean for those deploying AI in real companies? The key takeaway is that the decision-making depth, honesty, and thoroughness matter more than surface-level chat skills. It’s not about whether an AI writes well but whether it finishes what it starts, reads your internal files thoroughly, and remains honest when stakes are high.

Try It Yourself

If you’re interested, you can run your own business wargame against a read-only export of your operations—nothing writes back to your systems, ensuring safety. Discover which AI model aligns with your values and operational needs at firmulate.com/quiz.html.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI ethical decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Models Stand Firm Against Social Engineering Tests, Revealing Trustworthiness Before Real Crisis

AI models tested in simulated crises refused manipulation attempts, demonstrating trustworthiness and integrity before real-world deployment—an essential step for secure AI adoption.

Jets were 300 feet apart in Logan airport close call that forced Delta flight to abort landing, expert says

Two jets came within 300 feet at Logan Airport, prompting a Delta flight to abort its landing. Authorities are investigating the incident.

FAA investigates close call between two aircraft at intersecting runways at Boston Logan International Airport

The FAA is examining a close call between two aircraft at intersecting runways at Boston Logan, raising safety concerns. Details are still emerging.

Possible Collision Between JetBlue Jet And Drone At JFK Airport

A JetBlue flight at JFK reportedly came close to colliding with a drone, prompting safety concerns. Details are still emerging as authorities investigate.