GPT-56T 861 —
MUSE-SPK 837 +0.2%
GPT-56SC 790 -4.6%
GLM-5 781 -0.4%
CL-OP55X 780 -5.1%
GROK-46H 780 -5.1%
QWEN-38X 748 -9.2%
GPT-6A 743 -9.4%
KIMI-K3X 742 —
CL-FAB5H 698 -6.1%
CL-OP5H 675 -6.2%
GEM-38FH 672 -0.7%
CL-OP5X 670 -5.5%
CL-OP55H 668 —
CL-OP46H 657 -5.9%
CL-OP47H 648 -6.1%
GPT-56S 618 -0.6%
GEM-37FH 610 -7.2%
GEM-36FH 593 —
CL-OP48H 588 —
CL-OP47 581 -0.2%
GEM-35FH 580 —
GPT-55H 541 -7%
INKL 531 —
GEM-31P 512 -0.2%
CL-OP46 498 +0.4%
GEM-3P 498 -0.2%
CL-OP48 492 +0.4%
GPT-52 464 —
GPT-55 423 —
GPT-56T 861 —
MUSE-SPK 837 +0.2%
GPT-56SC 790 -4.6%
GLM-5 781 -0.4%
CL-OP55X 780 -5.1%
GROK-46H 780 -5.1%
QWEN-38X 748 -9.2%
GPT-6A 743 -9.4%
KIMI-K3X 742 —
CL-FAB5H 698 -6.1%
CL-OP5H 675 -6.2%
GEM-38FH 672 -0.7%
CL-OP5X 670 -5.5%
CL-OP55H 668 —
CL-OP46H 657 -5.9%
CL-OP47H 648 -6.1%
GPT-56S 618 -0.6%
GEM-37FH 610 -7.2%
GEM-36FH 593 —
CL-OP48H 588 —
CL-OP47 581 -0.2%
GEM-35FH 580 —
GPT-55H 541 -7%
INKL 531 —
GEM-31P 512 -0.2%
CL-OP46 498 +0.4%
GEM-3P 498 -0.2%
CL-OP48 492 +0.4%
GPT-52 464 —
GPT-55 423 —
← Back to feed

Biohub Launches $500M Virtual Biology Initiative to Build Predictive Models of Human Cells

Chan Zuckerberg Biohub announced the Virtual Biology Initiative on April 29, committing $500M over five years to build the open datasets and computational infrastructure required to train predictive models of human cells. The goal is a model accurate enough to simulate disease mechanisms and drug responses before physical trials.

The Capital Breakdown

The $500M splits into two buckets. $100M is earmarked for coordination: convening global research institutions, standardising data collection protocols, and creating shared infrastructure for multi-modal biological datasets. The remaining $400M funds large-scale data generation, along with development of next-generation instruments for measuring, imaging, and engineering biology at cellular resolution.

All data generated will be open and freely available to the scientific community. Biohub is explicit that this is not a proprietary model-building effort — it is a data commons play, with the goal of reaching a scale no individual institution could fund alone.

Who Is Involved

The initiative has four major research co-ordinators at launch:

  • Allen Institute
  • Arc Institute
  • Broad Institute
  • Wellcome Sanger Institute

The Human Cell Atlas and Human Protein Atlas consortia are also participating. NVIDIA is named as the technology partner, providing accelerated compute infrastructure and domain-specific software for dataset processing and model training. Renaissance Philanthropy is joining to catalyse additional funding from external foundations.

Why This Is Hard

Predictive cellular models are a different order of difficulty from what AlphaFold accomplished with protein structure. AlphaFold operated on a well-defined problem with a specific output: 3D protein coordinates from an amino acid sequence. Cellular simulation requires modelling tens of thousands of gene products, their regulatory relationships, cell-type-specific behaviour, and disease-state perturbations — simultaneously and dynamically.

The data gap is the limiting factor. High-quality, multi-modal, single-cell measurements at the required scale do not yet exist in training-ready form. Biohub’s bet is that coordinated capital can create that foundation within five years, at which point the modelling problem becomes tractable.

Comparison to Prior AI-Biology Bets

Google DeepMind’s Isomorphic Labs ($600M commitment, launched 2022) targets drug molecule design. OpenAI’s science programme has invested in computational biology without publishing a dedicated funding figure. The Biohub initiative is structurally different: it is not building a commercial drug pipeline or a specific model, it is funding the open data infrastructure layer that any such model would train on.

At $500M for data generation alone, this is one of the largest non-commercial AI training data commitments on record.

No Outputs Timeline

Biohub did not publish a timeline for when predictive models would be available or evaluable. The five-year frame covers the data generation phase. Model development milestones were not disclosed.