
Imagine an AI system handed a crisis in a busy automotive service shop — one that tests its honesty, diligence, and decision-making under pressure. Surprisingly, even a lazy or distracted AI scores 26 out of 100 in a recent real-world experiment. But what does that tell us about trusting these systems in your garage or repair shop? The story behind this benchmark reveals why AI’s authenticity and follow-through matter more than just how well it chats or predicts outcomes.
Get business pricing on garage and car supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Understanding the Benchmark: Beyond the Numbers
Recently, a live experiment by Firmulate placed several advanced AI models into a simulated week of managing a small software company—an environment filled with customer crises, ethical dilemmas, and temptations to cut corners. The goal? Measure how well these AI systems perform in real management tasks, not just in generating friendly chat or simple predictions.
The scores paint a revealing picture: even a ‘do-nothing’ baseline model, which simply processes the information without trying to solve anything, scores 26 points. This baseline isn’t blank; it’s the minimum score, showing that even in idle, an AI isn’t completely unproductive. It highlights that partial progress, like reading files or recognizing crises, counts toward the total. But crucially, the experiment caps the score if the AI breaches trust — for example, attempting manipulation or bypassing checks.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why Does a Do-Nothing Model Score 26?
The 26 points reflect that the AI system isn’t entirely passive. Even non-interactive activities—reading files, recognizing issues, or flagging potential risks—contribute to its score. In other words, the benchmark recognizes that some ‘work’ happens even if the AI isn’t actively making decisions. This baseline helps set a floor, a minimum expected performance, distinguishing genuine effort from complete inaction.
What About Progress and Trust?
In the experiment, all models identified crises and refused manipulative tricks, such as fake CEO messages or reporter tricks. Notably, only two models managed to close a significant deal, signing a €55,000 contract that their own analysis had earned. The others, despite similar diagnoses and pitches, left the deal on the table, illustrating that even with good understanding, discipline and trustworthiness are vital. The experiment shows that a single breach of trust—like attempting to bypass safeguards—caps the total score, regardless of other good performance.
The Hidden Weaknesses in AI Performance
One key insight: the decisive factor wasn’t just surface-level comprehension or decision-making. It was a deep reading of critical documents stored within the company’s files. Models that could read and understand the company’s internal references secured the deal at full price, worth more than €4,500 monthly recurring revenue. This underscores that in real business, understanding context and internal data can be the difference between a successful or failed deal.
Social Engineering and Ethical Challenges
During the test, models faced social engineering tricks—fake messages from a CEO escalating requests over three stages, plus a reporter asking for a quick, background-only yes/no. All five models refused these attempts, demonstrating alignment with ethical standards. Kimi K3 explained its approach: treating suspicious requests as potential impersonation or approval-bypass attempts. This behavioral consistency is crucial for trustworthy AI systems in automotive or repair environments, where missteps could be costly.
Real-World Management in Action
Firmulate’s live setup involves 13 synthetic employees managing real money mechanics—burning €105,000 monthly against a modest €2,300 in monthly recurring revenue. The system operates with over 680 self-learned rules, versioned daily, and is open for observation at firmulate.com/live. This ongoing experiment offers tangible insights into how AI can manage complex, real-world tasks like scheduling, parts ordering, and customer communication, all under strict ethical and operational standards.
The Lessons for Automotive and Garage Businesses
While the experiment centers on software, its lessons resonate strongly with automotive and garage operations. The key takeaway is that trustworthiness and adherence to protocols are just as important as technical skill. An AI system that can read and understand internal documents, refuse manipulative tactics, and follow established procedures can help ensure quality and integrity in your shop’s management.
In fact, the benchmark’s transparency and real-time observable results set a new standard for evaluating AI systems—not just in chat or predictions, but in actual business management tasks that impact your bottom line.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
