Basecamp Research has secured $140 million in funding from Nvidia, Anthropic's Anthology Fund, and other backers to build AI systems that learn from biological evolution. The London-based startup trains models on genetic sequences harvested from extreme environments like rainforests, ocean vents, and hot springs. These models then generate novel proteins, antibiotics, and therapeutics that don't exist in nature.

The approach inverts traditional drug discovery. Instead of screening millions of compounds in laboratories, Basecamp uses evolution as a training dataset. Nature has already run billions of experiments across millennia. The company extracts that signal from genomic data and teaches AI models to generate functional molecules.

Philip Lorenz, the company's CTO, positions biology as a harder problem than language for AI. Natural language processing benefits from massive text corpora and clear optimization signals: a sentence either parses correctly or it doesn't. Protein design lacks those scaffolds. A protein can fold in countless ways. Most folds are useless. Only rare configurations do useful work. The training signal is sparse and expensive to obtain through wet-lab validation.

This gap explains why language models scale predictably while biology models plateau. GPT systems improve reliably as you add data and compute. Protein models hit walls fast. Basecamp addresses this by starting upstream, at the genetic level. DNA sequences from real organisms encode evolutionary constraints that downstream protein structures must satisfy. By training on those sequences, the models learn which variations matter and which mutations break function.

The company's work spans two markets. Drug discovery focuses on antibiotic design, where bacterial resistance makes the need urgent and the chemistry is well-defined. Cell therapy tools aim at manufacturing optimized cells for cancer treatment and regenerative medicine. Both require designing molecules that work in living systems, not just on paper.

Basecamp's skepticism of benchmark scores reflects hard-earned biotech wisdom. A model can score perfectly on molecular property prediction tests and still generate proteins that aggregate, degrade, or fail to fold correctly. In silico validation tells you almost nothing about in vivo performance. The startup validates through experimental partnerships and wet-lab testing, slowing iteration but catching failures before expensive clinical trials begin.

The funding round signals investor confidence in biology-focused AI, but also caution. Nvidia and Anthropic don't fund undifferentiated machine learning anymore. They back startups with domain expertise, real data moats, and clear paths to revenue. Basecamp has all three. Its access to genetic databases from extreme environments creates competitive advantage. Environmental genomics remains fragmented and proprietary. A company that can license or partner for exclusive datasets gains asymmetric leverage.

The broader implication: AI in biology requires rethinking how we measure progress. Accuracy metrics borrowed from language modeling mislead. Biological AI succeeds when molecules fold properly, remain stable under physiological conditions, and produce measurable therapeutic effects. That demands integration with experimental validation, not just computational prediction.

Basecamp's bet on evolutionary data as training signal offers a clearer path than pure de novo design. The company sidesteps the need to fully understand protein biophysics from first principles. Instead, it learns empirically from nature's solutions.