
Imagine a company with no human employees, burning through €105,000 a month on a gamble of survival, and you can watch its every move unfold live. Welcome to the world of Firmulate, where artificial intelligence is running a small software business in a high-stakes test of trust, discipline, and decision-making, all in public view.
The Live Experiment: AI vs. Crisis
At the heart of this experiment is a tiny software company, staffed entirely by 13 synthetic employees powered by AI models. These models are given the same set of challenges—crises, customer requests, and ethical dilemmas—and must navigate them without human intervention. Every decision, every response, is versioned and recorded, creating a transparent, auditable record of AI behavior in a simulated real-world setting.
The Critical Test: The Same Week, Different Models
Four frontier AI models faced the same worst-week scenario: same customers, same crises, same temptations. The goal was to see which model could best manage the company’s operations, maintain discipline, and close a crucial €55,000 deal. The results? All four models identified every crisis and refused every manipulation attempt, showcasing strong ethical responses. Yet, only two of them succeeded in closing the deal based on their own analysis—highlighting a key insight: recognizing problems is not enough; executing the right decisions matters most.
The Hidden Weakness: Deadly Document References
Interestingly, the decisive advantage for the winning models was hidden in the company’s internal files. They read two document references deep into the files, uncovering a crucial piece of information that led to closing the deal at full price—an extra +€4,583 in Monthly Recurring Revenue (MRR). This demonstrates how reading and understanding context-rich data can be a game-changer in AI decision-making, especially when it comes to high-stakes negotiations.
Ethical Vigilance: Refusing to Be Fooled
The experiment also involved social engineering attempts. Fake CEO messages escalating in complexity, and even a reporter trick asking for a simple yes/no on background—each AI model refused these manipulative tactics. Kimi K3, one of the models, explained: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows that advanced AI can maintain integrity under pressure, refusing to be manipulated even when deception is sophisticated.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Inside the Company: A Real Money Machine in Real Trouble
Firmulate’s live company is a stark contrast to typical AI demos. It operates with 13 synthetic employees managing real money mechanics: burning €105,000 each month against a modest €2,300 monthly revenue. The company’s cash countdown is public, and its decisions are governed by over 680 self-learned playbook rules. Every day, the operation is versioned, and the system rebuilds itself twice daily at firmulate.com/live. This setup offers unparalleled transparency into how AI models perform under pressure in a real-world simulation.
The Profile of Opus 4.8
The most thorough model, Opus 4.8, with over 80 learned rules and deep analysis, finished last in the evaluation. It left a crucial deal on the table and slipped into internal discipline errors—writing attempts into a locked department instead of escalating. This highlights that even the most sophisticated models can falter under the weight of discipline and process management, not just technical capability.

The Scalpel and the Algorithm: Reclaiming Ethical Clarity and Clinical Confidence in the Age of AI
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why Should You Care?
The significance of this experiment isn’t just in AI running a mock company; it’s about what happens when AI takes on management roles, especially in critical business functions like CRM, support, or forecasting. The question is no longer whether AI can generate convincing chat; it’s whether AI can finish what it starts, read essential data, stay honest under pressure, and deliver measurable value.
Key Takeaways
- All models detected crises and refused manipulation, showing ethical robustness.
- The winning models applied deep reading, uncovering hidden data that led to closing deals at full price.
- Even with advanced discipline, AI can make process slips, underscoring the importance of process discipline and management.
- In this simulated environment, AI is tested as a real operational entity, not just a chatbot.

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Future of AI-Managed Business
This public, ongoing experiment makes clear that AI can engage in complex business scenarios with discipline and integrity. It can detect crises, resist manipulation, and make strategic decisions—though it still struggles with process discipline and execution. For entrepreneurs and managers considering AI as a decision partner, the message is simple: AI’s capabilities are advancing, but understanding its limits and ensuring disciplined processes remain crucial.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

SELL WITH AI: The Real Estate Agent's Playbook for ChatGPT, Claude, and AI Tools to Generate Leads, Write Listings, and Close More Deals
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.