The Data Moat Question

Sign in to track mastery·Sign in
competitive-strategydatamoats

The Data Moat Question

The central strategic debate in AI-bio: does proprietary biological data create a durable competitive moat, or are open-source foundation models and increasingly available public datasets eroding the advantage? The answer determines whether companies like Recursion (phenomics), Tempus (multimodal clinical), and Basecamp Research (biodiversity sequences) are building defensible franchises or expensive infrastructure that will be commoditized.

The bull case for data moats: Biological datasets are expensive and slow to generate. Recursion has >2.5 trillion cellular images across hundreds of disease contexts. Tempus has the largest private multimodal oncology dataset in the US (genomics, imaging, clinical records) — the AstraZeneca/Pathos deal (2024) and consistent ~126% net revenue retention (2025) suggest customers can't easily replicate it. Basecamp Research has sequenced organisms from extreme environments that no one else has sampled. Unlike software, you can't easily scrape or reverse-engineer a proprietary cell-painting dataset.

The bear case: Foundation models like ESM3 and the PDB (Protein Data Bank, now >200,000 structures) are public. Any company with sufficient compute can train a competitive protein model without proprietary data. Schrodinger's physics-based FEP+ is differentiated by technique, not data. And open-source models from Meta, EvolutionaryScale, and Baker Lab are eroding the moat of closed platforms.

The nuanced reality in 2026: The moat question is molecule-type-specific. For small-molecule screening, open data (ChEMBL, PubChem) + public models have substantially closed the gap. For phenomics / cell morphology (Recursion's core), proprietary scale still matters — Recursion's 2.5T images are 10-100x any academic equivalent. For multimodal clinical data (Tempus), HIPAA constraints protect the moat from being replicated quickly. For protein foundation models, ESM3 being open-weight means the base model moat is largely competed away; the application layer moat (domain fine-tuning, closed-loop integration) remains.

Key Claims

  • Recursion has the largest phenomics dataset — >2.5T cellular images — Built over 10+ years; used for disease biology mapping and target identification. Evidence: strong (Recursion)
  • Tempus's data moat is validated by 126% NRR (2025) — Customers expand their data spend once they integrate; strong signal of lock-in. Evidence: strong (Tempus SEC 8-K FY2026, theaiinsider.tech)
  • Open protein models are commoditizing the base model layer — ESM3 (98B params) is open-weight; RFdiffusion and ProteinMPNN are MIT-licensed. Closed platform moats on foundation models are eroding. Evidence: strong (EvolutionaryScale)
  • Basecamp Research's biodiversity data is genuinely novel — Samples from extreme environments not in any public database; key for novel enzyme discovery; private and funded for expansion. Evidence: moderate
  • Proprietary data alone is insufficient without the assay-to-model feedback loop — The companies winning are those that generate new proprietary data continuously (Recursion's automated labs), not those sitting on a static dataset. Evidence: moderate (morningglorysciences.com)
  • Sequence-scale is now a generated asset, not just a scraped one — Biohub's ESM Atlas spans 6.8B sequences and 1.1B predicted structures — The former-EvolutionaryScale team trained ESMC on ~2.8B sequences across the tree of life and generated the ESM Atlas (6.8B sequences, 1.1B predicted structures) "in a couple of weeks." This reframes the protein-data moat: predicted-structure scale can be manufactured rapidly with compute, reinforcing the bear case that base-model/structure data commoditizes — while the wet-lab-validated binder results (against five named disease targets) are the harder-to-replicate output. Evidence: moderate (Biohub world model)

Open Questions

  • Does phenomics data advantage translate into better clinical outcomes, or only into better target identification?
  • Will public health systems (NHS, Kaiser, VA) create public multimodal clinical datasets that erode Tempus's advantage?
  • Is Basecamp Research's biodiversity moat defensible or will the diversity of public databases (UniProt, MGnify) catch up?

Related Concepts

Changelog

  • 2026-06-15 — Initial compilation; Recursion phenomics, Tempus NRR, ESM3 open-weight dynamic covered
  • 2026-06-24 — Compiled new sources (biohub-protein-world-model)

Theses that depend on this concept

These research positions cite this concept in their evidence. If the concept changes materially, these theses may need re-scoring.

The Data Moat Question | KB | MenFem