Andrej Karpathy, OpenAI co-founder and former VP of AI, has launched a quest to define the next benchmark for AI capability. His method: vibe tests that push models beyond standard benchmarks.
Karpathy fed Claude Opus 5 a single paragraph from J.R.R. Tolkien's "Lord of the Rings" prologue. The model generated 5,500 lines of code that rendered Tolkien's opening scene as an interactive 3D browser experience. The test wasn't about accuracy or speed. It measured something harder to quantify: whether the AI could translate literary prose into functional, creative output that captures the writer's intent.
This approach reflects a growing frustration in AI circles with narrow performance metrics. Standard benchmarks test factual recall, math skills, and coding ability in isolation. They miss what Karpathy calls "vibe"—the nebulous quality of whether an AI actually understands context, tone, and nuance well enough to apply them creatively.
The Lord of the Rings test exemplifies this thinking. Claude didn't just summarize or analyze Tolkien. It understood the medieval fantasy atmosphere, the sense of scale and wonder, and the need for spatial visualization. Then it executed that understanding through working code.
Karpathy's search for new vibe tests reflects a shift in how AI builders think about progress. As models hit plateaus on traditional benchmarks, engineers look for harder, more subjective measures. Can an AI generate genuinely novel ideas? Can it collaborate with humans on open-ended creative projects? Can it handle ambiguity and produce outputs that feel "right" rather than merely correct?
This matters because benchmarks drive development priorities. If the field optimizes only for test scores, it risks building systems that excel at pattern matching while struggling with real-world reasoning. Karpathy's emphasis