85% of enterprises that experienced a test-passing AI agent fail in front of customers are now pursuing fully automated, no-human-approval deployment, compared to 61% of companies with no such incident, according to VentureBeat Pulse research across 108 enterprise respondents in July. Trust in automated evaluation rose from 5% to 13% month-over-month. Customer-visible failure rates held flat: 49% in July, 50% in June, across 265 combined responses.

The data contains a sixfold confidence gap that tells the real story. Among companies that had never experienced a testing miss, 24% expressed complete faith in automated checks. Among those that had been burned, only 4% did. Yet the burned group is removing humans from deployment pipelines at a higher rate anyway. Raindrop CTO Ben Hylak told VentureBeat that Fortune 100 companies are actively shrinking eval sets and deprioritizing maintenance, shifting instead toward anomaly detection in production as agent systems grow too complex for enumerated test cases.

The full report is worth reading for the mechanics of what it calls the evaluation gap: the split between release-gate confidence and actual production reliability, the shift in buying criteria toward integration ease, and the Braintrust platform share gain. The unresolved question the data raises is this: if per-deployment failure rates stay constant while deployment volume scales and human review disappears, the total incident count rises even if the percentage of affected companies does not. July did not measure incident volume. That number is next.

[READ ORIGINAL →]