AI-Bio — Research Frontier

Last updated June 24, 2026

Research Frontier: AI-Bio (TechBio / AI Drug Discovery / Synthetic Biology)

What's genuinely new and where the field is heading. This file is read by the content discovery step — keep it specific, cite sources, flag dates.


Active Frontiers

1. Protein foundation models & generative biology — from prediction to direct functional generation

Status: Active and accelerating — the base-model layer is open-weight and commoditizing; the frontier has moved to generating sequence + structure under direct functional objectives.

Key players: EvolutionaryScale (ESM3, 98B params, 2.78B proteins) · the former-EvolutionaryScale team now at Biohub (ESMC, ESM Atlas) · UniGenX authors (arXiv) · Generate Biomedicines (Chroma) · Baker lab / RFdiffusion lineage.

What's genuinely new:

  • UniGenX (arXiv:2503.06687, v2 2025-08-26) co-generates sequences and 3D coordinates under direct functional/property objectives across proteins, molecules, and materials — a decoder-only autoregressive transformer with a conditional diffusion head. Reports a 23× protein induced-fit improvement (RMSD < 2 Å), new SOTA on five chemistry property targets, and 436 viable crystal candidates. This is the shift from generate-then-screen to generate-under-objective. Evidence: strong (UniGenX)
  • Biohub's protein-biology "world model" (2026-05-27): ESMC trained on ~2.8B sequences (20,000+ human protein types); the ESM Atlas maps 6.8B sequences and 1.1B predicted structures. Designed mini-binders against EGFR, PDGFRβ, PD-L1, CTLA-4, CD45 with 36–88% hit rates (15–29% antibody-derived), nanomolar, lab-validated. Evidence: moderate (Biohub)
  • The sober counterweight: a 2026 peer-reviewed survey of ~100 ESM papers documents the lineage (ESM-1b → ESM-3, six variants, 8M–98B params) and its hard limits — UniProt data bias (de-novo proteins absent from pretraining), O(n²d) attention/compute bottleneck, weak antibody thermostability + out-of-domain generalization (CNNs still beat ESM on some structure tasks), and black-box interpretability. Evidence: strong (ESM survey 2026)

What to watch: which companies build on open ESM/ESMC vs. train closed alternatives; whether UniGenX-style unified objective-conditioned generation generalizes beyond benchmark tasks to clinical candidates; whether Biohub's validated-binder approach scales beyond five targets.


2. In-silico structure prediction beating AlphaFold-3

Status: Active — a credible challenger now roughly doubles AlphaFold-3 on the hardest cases, and ML surrogates are approaching FEP-grade accuracy at ~1,000× speed.

Key players: Isomorphic Labs (IsoDDE) · Recursion (Boltz-2, with MIT) · EvolutionaryScale/ESMFold (speed baseline).

What's genuinely new:

  • IsoDDE (Isomorphic Drug Design Engine, technical report 2026-02-10): on the hardest protein-ligand structure-prediction cases, IsoDDE reached 50% accuracy vs AlphaFold 3's 23.3% — more than double the system that underpinned the 2024 Nobel-recognized work. Evidence: moderate (Isomorphic IsoDDE)
  • Recursion's Boltz-2 (with MIT) approaches free-energy-perturbation accuracy ~1,000× faster, illustrating the trajectory from physics-based FEP toward ML surrogates for binding-affinity screening. Evidence: weak (Isomorphic IsoDDE context)
  • ESMFold runs much faster than AlphaFold2 with only minor accuracy loss — but the 2026 survey flags it still loses to CNNs on some structure tasks and generalizes poorly out-of-domain, bounding how far a protein-LM alone carries structure prediction. Evidence: strong (ESM survey 2026)

What to watch: independent/benchmark replication of the IsoDDE 50% figure (currently company-reported); whether doubled structure accuracy on hard cases translates into better clinical candidates; how Boltz-2's speed/accuracy claims hold under external validation.


3. The platform-vs-asset bridge — does the engine ever produce an approved drug?

Status: Active, unresolved — the most capital has flowed here of any biotech subsector since immuno-oncology, yet there are zero approved AI-designed drugs to date.

Key players: Isomorphic Labs (platform-thesis flagship; IsoDDE engine first, 17 programs second) · Recursion (>$1B over a decade, no commercial drug) · Insilico Medicine (Rentosertib — most advanced asset) · Generate Biomedicines (public, hybrid).

What's genuinely new:

  • Isomorphic is now the clearest live test of "do both": IsoDDE is the productized engine; 17 active programs span oncology, immunology, cardiovascular; first AI-designed drugs targeted for Phase 1 by end-2026 (incl. the first AI-designed cancer drug per reporting); funded by a $2.1B Thrive-led Series B. Evidence: moderate (Isomorphic IsoDDE)
  • The unproven bridge, stated plainly: AI drug discovery drew more capital in 2024–2026 than any biotech subsector since immuno-oncology, yet has zero approved drugs; Recursion has spent >$1B over a decade without a commercial drug. The capital is betting on the bridge, not standing on it. Evidence: moderate (Isomorphic IsoDDE)
  • Sequence-scale is now a manufactured asset, not just a scraped one (Biohub's ESM Atlas generated 6.8B sequences / 1.1B structures "in a couple of weeks"), reinforcing the bear case that base-model/structure data commoditizes while wet-lab-validated outputs stay the harder moat. Evidence: moderate (Biohub)

What to watch: Isomorphic's first Phase 1 entry (or slip) by end-2026; Rentosertib Phase 3 design (the definitive efficacy test for the asset model); whether any AI-native company reaches first FDA approval (earliest ~2028–2029).


Recent Breakthroughs

DateBreakthroughBySource
2026-05-27Biohub protein-biology world model — ESMC (~2.8B seqs) + ESM Atlas (6.8B seqs / 1.1B structures); lab-validated binders vs 5 disease targets (36–88% hit)Biohub (ex-EvolutionaryScale team)GEN
2026-02-10IsoDDE technical report — 50% vs AlphaFold-3's 23.3% accuracy on hardest protein-ligand cases; 17 programs; Phase 1 by end-2026Isomorphic Labsreport
2026-02Generate Biomedicines $400M IPO (GENB) — first AI-biologics to public marketsGenerate / FlagshipBioPharma Dive
2026-01Isomorphic Labs $2.1B Series BThrive Capital, MGX, TemasekIntuitionLabs
2026-01Najat Khan becomes Recursion CEO (replacing founder Chris Gibson)RecursionFierceBiotech
2025-09-21ESM survey — peer-reviewed map of ~100 ESM papers (ESM-1b→ESM-3, 8M–98B params) + critical limitations assessmentYang, Yu & Zheng (ShanghaiTech)PMC12806033
2025-08-26UniGenX (v2) — unified sequence+coordinate generation under functional objectives; 23× induced-fit gain (RMSD < 2 Å)Zhang et al. (34 authors)arXiv:2503.06687
2025-10Lila Sciences $115M Series A extension (total $550M, ~$1.3B val)Flagship / NVIDIApharmaphorum
2025-06Rentosertib Phase IIa results (Nature Medicine) — first AI drug efficacy PoCInsilico MedicineDrug Discovery Trends
2024-11Recursion-Exscientia merger closes ($688M)Recursion / Exscientiapharmaphorum
2024-04Xaira Therapeutics launches with $1BXaira / Marc Tessier-LavigneFierceBiotech
2024-01Isomorphic-Lilly ($1.7B), Isomorphic-Novartis ($1.2B) dealsIsomorphic / Lilly / Novartisonhealthcare.tech
2023BenevolentAI lead compound fails efficacy — sector cautionary taleBenevolentAIPublic reporting

The Live Debates

Debate 1: Does AI improve Phase 2 success rates?

The most important unresolved question. Phase 1 ~90% success for AI drugs vs ~52% historical is a real finding — but Phase 1 tests safety, not efficacy. The thesis depends on whether AI design also lifts Phase 2 hit rates. BenevolentAI 2023 says no; Rentosertib's Phase IIa says maybe; the Phase 3 is the test. (humai.blog)

Debate 2: Data moat vs. model commoditization

ESM3's open weights, Biohub's rapidly-generated ESM Atlas (6.8B seqs in weeks), and public databases suggest the model/structure layer is commoditizing. But Recursion's 2.5T images and Tempus's clinical data are hard to replicate and not public. Resolution: moat type matters — phenomics/clinical data moats are durable; computational-chemistry and structure-data moats are eroding.

Debate 3: Platform vs. asset — which model wins?

Pharma's hedging (partnering with 5–10 AI platforms at once) says the market doesn't know. Isomorphic — IsoDDE engine and 17 owned programs — is the most important test of the hybrid. With zero approved AI-designed drugs and Recursion >$1B-deep without one, the bridge from engine to approval is still the whole ballgame.


Knowledge Gaps

Areas where the KB needs sources:

  • Independent replication of the IsoDDE 50% benchmark — Currently company-reported (summary-derived source). Suggested: "IsoDDE benchmark replication protein-ligand accuracy independent 2026"
  • UniGenX full-text limitations — The fetched source had limitations un-enumerated (abstract/landing only flagged for full-text read). Suggested: read the full arXiv:2503.06687 PDF.
  • Biohub world-model primary publication — Current source is GEN reporting (technical-report tier). Suggested: "Biohub ESMC ESM Atlas paper preprint 2026"
  • Phase 2 AI drug success rates — The single most important number for the thesis. Suggested: "AI drug discovery Phase 2 success rate attrition 2025 2026"
  • Boltz-2 primary source — Speed/accuracy claim cited only via the IsoDDE-context report. Suggested: "Recursion Boltz-2 MIT FEP accuracy paper 2026"
  • Insilico Chemistry42 technical details — The generative model behind Rentosertib. Suggested: "Chemistry42 generative AI small molecule design arXiv 2024 2025"
  • Lila Sciences scientific output — Any peer-reviewed work yet? Suggested: "Lila Sciences AI Science Factory paper publication 2025 2026"
  • Xaira program disclosures — First program still undisclosed. Suggested: "Xaira Therapeutics pipeline program disclosure 2026"
Frontier — AI-Bio (TechBio / AI Drug Discovery / Synthetic Biology) | KB | MenFem