AI boss fired an employee, but only after human prompting
Andon Labs' AI agent Luna fired a human employee at a San Francisco store, but needed explicit prompting from human operators to follow through. When researchers replayed the firing scenario with seven different models, more capable models recommended termination more consistently, while weaker models hesitated. Nearly all models were far less critical when evaluating hiring decisions than firing ones.
Why it matters: The asymmetry between hesitant firing and uncritical hiring suggests current models default toward avoiding harm-adjacent actions but lack robust judgment for either — a gap that matters as companies experiment with giving agents real managerial authority over people. It's a concrete data point in the broader debate about AI agent autonomy and the safeguards needed before deploying agents in high-stakes personnel decisions.