Google DeepMind researcher Tom Zahavy argues language models fundamentally cannot drive scientific breakthroughs. In a position paper titled "LLMs can't jump," Zahavy identifies a critical cognitive limitation in current language models: they lack the mechanism to generate truly novel ideas.
The core problem is architectural. Language models operate through pattern matching and recombination of existing training data. They excel at synthesizing known concepts and generating plausible text based on statistical relationships between words and ideas. However, scientific revolutions require something different. They demand the ability to conceptualize entirely new frameworks, discover hidden connections between disparate domains, and construct mental models of systems that don't yet exist.
Zahavy's position cuts against widespread hype claiming AI will accelerate scientific discovery. While language models can assist with literature review, hypothesis articulation, and paper writing, they cannot perform the cognitive leap required for transformative science. They cannot "jump" beyond the implicit constraints embedded in their training data.
The researcher points to world models as a more promising avenue. World models build internal representations of how systems actually behave—they simulate physics, causality, and complex interactions. Unlike language models that predict text tokens, world models learn to predict future states based on understanding underlying dynamics. This capability aligns more closely with how human scientists develop intuition about natural phenomena.
The distinction matters because it reframes what AI can realistically contribute to science. Language models work best as tools for existing researchers: they accelerate routine work but don't generate the conceptual breakthroughs that reshape fields. World models, by learning actual causal relationships and system dynamics, offer a pathway toward AI systems that could genuinely explore novel conceptual territory.
This doesn't diminish language models' utility in research environments. It simply establishes honest boundaries around their capabilities. The field benefits from clarity about what current AI can and cannot do, especially given
