Trust in automated agent evaluation jumped sharply among enterprises in July, but the underlying problem persisted. Across 108 companies surveyed, organizations that rely on automated testing to validate AI agents nearly tripled their confidence in these systems, rising from 5% to 13%. Complaints about evals failing to predict real-world performance dropped 10 points.

Yet the data reveals a dangerous disconnect. Nearly half of all surveyed enterprises shipped an agent that passed its internal evaluations and then failed in production. The trust increase came almost entirely from companies that have never experienced this failure. Among organizations that got burned, only 4% trust automated evaluation. Among those that haven't, 24% do.

The counterintuitive finding: getting burned accelerates the removal of humans from the loop, not slows it. Enterprises that suffered agent failures don't respond by tightening oversight. Instead, they double down on automation and reduce human review.

This pattern suggests a fundamental misalignment between how companies test agents and how those agents perform in the wild. Automated evaluation systems pass agents that later fail customers. Rather than diagnosing why evaluations miss real-world edge cases, burned organizations appear to interpret the failure as a cost of moving faster. They pull humans out of the decision loop to reduce the friction they perceive as slowing deployment.

The data comes from the second wave of ongoing research tracking enterprise adoption of autonomous agents. It reveals a critical gap in how organizations validate mission-critical automation. Companies building confidence in faulty evals, and companies that have been proven wrong by the real world responding not with caution but with accelerated autonomy, creates systemic risk.

For vendors selling evaluation tools, the implication is stark: your customers don't trust you after they discover your product failed them. For enterprises deploying agents, the lesson is darker: the faster you move toward full autonomy, the less visibility