Across 108 enterprises surveyed in July 2026, full trust in automated agent evaluation nearly tripled from 5% to 13%, and the failure rate that trust is supposed to reflect did not move one point. Just under half of organizations (49%) shipped an agent that passed internal evals and then caused a customer-facing failure, statistically identical to June's 50%. Trust improved. Correctness did not.

The cross-tabs are where the story lives. Among enterprises that have experienced a false-confidence failure, only 4% fully trust automated evaluation. Among those that have not been burned yet, 24% do. The new confidence belongs almost entirely to the inexperienced. Worse, the burned enterprises are not pulling back from autonomy: 85% of them already allow zero-human deployment or are engineering toward it, against 61% of those with no failure history. Getting burned accelerates the march toward removing humans from the loop, not slowing it.

On the tooling side, the market is consolidating. Enterprises running no dedicated evaluation tooling dropped from 17% to 12%. Braintrust nearly doubled to 15% and DeepEval reached 17%. Switching intent cooled, with those planning no change rising from 36% to 44%, and ease of integration overtook cost as the top selection factor, jumping from 27% to 39%. The full report, part of VentureBeat's Pulse Research series, is worth reading for the methodology section alone: the researchers flag a composition shift between waves, with Technology and Software firms dropping from 23% to 14% of respondents, which means some of the trust increase may reflect who answered rather than what changed.

[READ ORIGINAL →]