
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
What if your next CEO was an AI — and it was struggling to keep the lights on?
Imagine a startup where every decision, crisis, and opportunity is made by artificial intelligence — and you can watch it unfold live, every workday. This is no science fiction; it’s the core idea behind Firmulate, a public experiment in AI management that reveals crucial insights into how machines handle real-world business challenges.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Live Company: A New Way to Test AI Performance
Firmulate operates a unique, publicly visible small software company, driven entirely by AI models. This company runs daily, facing the same customer demands, crises, and temptations that any real business would encounter. What makes it extraordinary is that every decision is versioned, auditable, and transparent — giving followers a front-row seat to how AI handles complex management tasks.
Since its launch, the company has been under constant watch. It burns €105,000 each month against a revenue of only €2,300, with a public cash countdown highlighting its fragile state. Each workday, the company updates with new decisions, new rules, and new challenges, all available for real-time viewing at firmulate.com/live.
Measuring AI’s Strategic Acumen
Pressed to handle crises and negotiate deals, four frontier AI models were tested against the same set of challenges. They faced customer crises, manipulation attempts, and internal ethical dilemmas. Remarkably, all four models identified every crisis and refused every manipulation — including a staged social engineering attack involving fake CEO messages. When it came to closing a crucial deal, only two models signed the €55,000 contract their own analysis justified, highlighting a key difference: the ability to read deeper into the company’s own files for hidden opportunities.
Hidden Opportunities and Critical Failures
The decisive advantage came from models that delved into internal documents, uncovering a buried fact that boosted the company’s revenue potential by over €4,500 monthly. This demonstrates that in real-world settings, the deepest understanding often comes from thorough reading and analysis, not surface-level conversation.
Building in Public — With Real Money and Real Risks
The entire experiment is about more than just AI scores. It’s a window into how these models perform under pressure, how disciplined they stay, and whether they heed organizational rules. The company’s 13 synthetic employees, guided by over 680 self-learned rules, are not just bots but actors in a real economic system, risking real money and facing real deadlines.
For instance, when approached with a staged social engineering scam, all models refused to cooperate, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows a level of discipline that chat demos rarely reveal.
Implications for Business and AI Adoption
This experiment raises critical questions for any organization considering AI integration. It’s not sufficient for AI to generate convincing chat or support responses; it must also perform reliably in decision-making, read internal data thoroughly, and resist manipulation. In the current setup, only half the models secured the deal based on their own analysis, emphasizing that performance under pressure is what truly counts.
The League of AI Performers
Among the models tested, GPT-5.6-SOL led with a score of 95, demonstrating it could find the buried fact and close the deal. Kimi K3 followed closely at 93, with Sonnet 5 and Opus 4.8 trailing behind. The performance scores reflect their ability not just to identify crises but to act decisively and ethically — critical factors for real-world deployment.
Beyond the Demos: Real Business, Real Stakes
The live site offers a transparent view into a working company that is actively losing money, yet providing invaluable data on AI management capabilities. Business leaders, AI developers, and strategists can watch as decisions unfold, analyze what makes some models succeed where others falter, and consider how these insights might translate into their own operations.
Whether you’re a student, researcher, or executive, this experiment is a rare chance to see AI in the trenches — making decisions that matter, under real constraints, in real time. It’s a glimpse into the future of autonomous management, and the lessons learned could shape how we build AI systems that truly work for us.

Key Takeaway
Watching this AI-managed company in action reveals that success depends not just on language skills but on the ability to thoroughly analyze, resist manipulation, and make decisive, trustworthy choices under pressure. The experiment offers a rare, transparent look at AI’s potential and limitations in real-world management.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Pool season Picks
robotic pool cleaners
As an affiliate, we earn on qualifying purchases.