The training data that fueled large language model development faces a looming scarcity problem. The internet's clean, freely available text that powered the first wave of LLM advances is becoming increasingly contaminated with AI-generated content, subject to legal disputes over ownership rights, and expensive to replace with fresh data.

AI companies are responding by acquiring older materials. Reports emerged this week of AI firms purchasing historical book collections to supplement their training datasets. This shift signals recognition that the easy phase of data collection has ended. Mining the public web for text now yields diminishing returns as AI-generated content proliferates across platforms, making it harder to distinguish human-created material from machine-generated output.

The data scarcity challenge extends beyond text. Companies face two compounding obstacles. First, content creators increasingly contest AI's use of their work in training, with authors and publishers filing lawsuits and implementing technical barriers. Second, sourcing genuinely novel, high-quality data at scale requires spending capital that early-stage LLM development avoided.

Nvidia tackled a related problem on the robotics front this week by releasing a simulator that trains robots through video, motion, and synthetic consequences. Rather than waiting for real-world data from actual robot deployments, this approach generates training examples computationally. Synthetic data becomes a workaround to the data scarcity problem, though quality and transferability to physical systems remain open questions.

The broader implication: AI companies cannot simply scale existing approaches indefinitely. The era of training models on cheap, permissionless internet content is ending. Future development requires either spending on data acquisition, licensing, or relying on synthetic alternatives. This transition reshapes the economics of AI development and raises questions about which companies possess the capital and legal frameworks to sustain training at scale.