Individual AI agent conversations can score perfectly while revealing fundamental product failures. This gap is reshaping how enterprises evaluate AI systems, pushing evaluation methods away from scoring isolated interactions toward analyzing entire user cohorts against performance baselines.

Harrison Chase, CEO of LangChain; Hui Zhang, CTO and co-founder of Conviva; and Emmanuel Turlay, director of engineering at CoreWeave discussed this shift at VB Transform 2026. The change reflects a maturing market recognizing that single successful exchanges mask systemic problems across deployed agents.

The industry is simultaneously moving toward cheaper, narrower judge models for evaluation. Agent-as-judge systems, where one AI agent evaluates another's output, haven't replaced LLM-as-judge approaches, which remain the default evaluation method. Chase confirmed this standard persists despite emerging alternatives.

The broader tension Zhang highlighted centers on automated judging mechanisms versus other evaluation approaches. Traditional single-trace scoring creates false confidence. A chatbot interaction might execute flawlessly while the underlying system fails to handle edge cases, context switching, or consistency across multiple conversations. Cohort analysis reveals these patterns by comparing how user groups experience the product under real conditions, not just cherry-picked examples.

This reflects lessons learned as enterprises deployed agents in production. Individual conversation quality doesn't predict overall reliability. A system might handle one customer query perfectly but consistently stumble on similar requests from different users or fail when context changes.

The movement toward cheaper judge models addresses a practical concern. Evaluating agents at scale demands efficient tooling. Narrower, purpose-built models reduce computational costs while maintaining accuracy for specific evaluation tasks. This approach balances thoroughness with economic viability as enterprises scale agent deployments.

The shift signals maturation in AI evaluation practices. Early adoption focused on demonstrating individual capabilities. Production deployment demands systematic, cohort-level understanding of reliability and consistency. Enterprises now recognize that