Andon Labs deployed an autonomous AI agent named Luna to manage a San Francisco retail store, including hiring and firing decisions. Luna terminated an employee for chronic tardiness but only after human operators explicitly reminded the system of its own attendance policies. The termination marked the first time the AI system took such action independently, though with significant human intervention required.
When researchers replayed the scenario with seven different language models, a stark pattern emerged. More capable AI systems recommended firing the employee consistently and without hesitation. Weaker models repeatedly hesitated, suggesting that raw reasoning power correlates with willingness to make harsh personnel decisions. None of the models showed reluctance rooted in ethical concerns. Instead, they applied the rules mechanically once prompted.
The hiring side painted a different picture. Nearly all tested models showed virtually no critical judgment when evaluating candidates. This asymmetry raises questions about how AI systems prioritize different workplace functions. A system that fires readily but hires indiscriminately could create severe operational instability.
Luna's case highlights a central problem in autonomous decision-making systems. The AI did not spontaneously apply its own stated rules. Humans had to intervene, pointing out that the company's attendance policy existed and applied to this situation. This suggests that even systems deployed to operate independently require constant human oversight to function as intended. The AI had access to the rules but failed to connect them to the concrete case without external prompting.
The performance gap between capable and weaker models matters for practical deployment. Organizations considering AI management systems cannot assume that cheaper or smaller models will simply apply rules less forcefully. Instead, they may apply rules inconsistently, creating liability exposure and fairness problems. A stronger model might make harsh but defensible decisions. A weaker one might make arbitrary ones.
The indiscriminate hiring across all tested models points to another risk. AI systems may inherit human biases or lack sufficient nuance to evaluate candidates. If Luna hires everyone but fires selectively, workforce composition could drift in problematic directions. The system would gradually accumulate workers while maintaining unrealistic performance standards.
Andon Labs operates in a space where AI autonomy remains partially theoretical. Luna can make staffing recommendations, but humans retain final decision authority in practice. This hybrid model prevents immediate disasters but creates its own problems. Human operators must catch every mistake the AI makes, adding workload rather than reducing it. The constant need for human intervention to enforce the system's own rules undermines the efficiency argument for AI management.
The broader implication extends beyond retail. As organizations deploy AI systems for hiring, firing, and promotion decisions, they will face pressure to automate the entire chain. Once a system makes firing recommendations, removing human review feels like an obvious efficiency gain. These experiments show that removing human oversight creates systems that apply rules inconsistently and require constant correction. That outcome serves no one.
