Enterprise AI teams are deploying autonomous agents to production with broken safety guardrails. Half of 157 surveyed enterprises have already shipped an agent that passed internal tests but failed in the real world. Yet two-thirds are moving toward fully automated deployment decisions with no human oversight.
The core problem is not a lack of testing. Organizations have evaluation frameworks. The problem is that these evaluations do not predict real-world performance. Only 5 percent of enterprise teams fully trust their automated evaluation systems. Most cite a single, damning weakness: their tests do not align with actual customer outcomes.
This creates what researchers call an "evaluation gap." Companies grant agents increasing autonomy while losing confidence in the mechanisms that are supposed to prevent failures. The tension is acute. Faster deployment requires removing humans from the loop. But removing humans requires trustworthy evaluations. Trustworthy evaluations require tests that actually predict production behavior. Few enterprises have built those tests.
The data shows the gap widening. Two-thirds of surveyed organizations are either already using automated evaluation to gate production deployments or actively engineering systems to do so. They are moving faster, not slower, despite knowing their evaluations are unreliable. The incentive structure rewards speed. The cost of a botched autonomous agent in production appears, in many cases, to be treated as acceptable.
This is not a technical problem with easy fixes. Simulating real-world complexity in a test environment is hard. Customer behavior is messy and context-dependent. The agents these enterprises are deploying handle variable requests across domains where edge cases are common. Building evaluations that catch failures before they reach customers requires understanding what failures actually look like in production. Most enterprises have not done that work yet.
The research suggests enterprises are racing ahead of their own confidence. They know their evaluation systems are weak. They know half have already burned customers. They are deploying faster anyway. This pattern
