
Imagine a team of AI managers running a real software company, facing the same crises, temptations, and deadlines. Would they deliver consistent results? Or would their personalities make all the difference? In a groundbreaking live experiment, four frontier AI models have been put to the test in a simulated company environment—revealing not just their decision-making skills, but their management personalities and ethical fibers. If you’re invested in AI’s role in business, understanding these differences could reshape how you evaluate and deploy these digital managers.
Prime for Young Adults — start your free trial
Fast free delivery, streaming and member deals for eligible 18–24 year olds.
As an affiliate, we earn on qualifying purchases.
The Real-World Experiment: Putting AI to the Test
At the heart of this experiment is a small, real software company running every business day with actual money mechanics: a monthly burn rate of €105,000 against a revenue of just €2,300, a public cash countdown, and over 680 self-learned playbook rules. Each day, the AI models are tasked with managing crises—customer complaints, security breaches, ethical dilemmas—just like human managers. The decisions they make are fully versioned and auditable, ensuring transparency and the ability to analyze their behaviors.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Models and Their Scores
- GPT-5.6-SOL: Scored the highest at 95, successfully identifying critical information buried two documents deep in the company’s files, and closing the key €55,000 deal.
- Kimi K3: Slightly behind with a score of 93, Kimi signed the deal, demonstrating the cleanest discipline and a keen eye for detail, despite running without an effort parameter (default API settings).
- Sonnet 5: Scored 88, managed to close the deal but with some process slips.
- Fable 5: Scored 77, also closed the deal but showed weaker process discipline and left some opportunities on the table.
- Opus 4.8: Last place at 73, with discipline slipping and decisions left unresolved, despite thorough analyses with over 80 learned rules.
AI ethics and crisis management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Sets These Models Apart?
All models excelled at crisis detection and refused manipulation attempts, including social engineering scenarios like staged CEO messages and a fake reporter request—demonstrating ethical acuity. Yet, the decisive factor was their ability to read and act on information buried deep within company files. Only the top two models identified and used this crucial data to secure the deal, resulting in a significant revenue boost of over €4,500 per month.
AI business performance monitoring tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Management Personalities in AI
The experiment revealed that AI models exhibit management personalities akin to human traits: some are meticulous and disciplined, others are more opportunistic or hesitant. For example, Opus 4.8, despite being detailed and analytical, left opportunities on the table due to slips in process discipline. Kimi K3 maintained a straightforward, disciplined approach, which contributed to its success. Notably, the default API settings used by Kimi may have contributed to its clean decision-making, highlighting how configuration influences management style.
enterprise AI management solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why It Matters for Your Business
If AI tools are to support your CRM, handle customer support, or forecast sales, their ability to stay honest, read critical information thoroughly, and follow disciplined processes matters more than their conversational skills. The real-world performance—deciding to sign or reject a deal, refusing social engineering—provides a clearer picture of their potential than chat demos ever could.
Experience the Live Performance
Curious? Watch the company’s day-to-day operations in real-time at firmulate.com/live. See the decisions unfold, get insights from actual employee messages, and test your intuition by guessing which AI model made each choice at firmulate.com/quiz.html.

The live experiment demonstrates that AI management personalities matter: some models are disciplined and thorough, others opportunistic or slip-prone. For your business, choosing an AI with a proven track record of integrity and thoroughness isn’t just smart—it’s essential for trustworthy automation.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.