Biohub Launches $500M Virtual Biology Initiative to Build Predictive Models of Human Cells
Chan Zuckerberg Biohub announced the Virtual Biology Initiative on April 29, committing $500M over five years to build the open datasets and computational infrastructure required to train predictive models of human cells. The goal is a model accurate enough to simulate disease mechanisms and drug responses before physical trials.
The Capital Breakdown
The $500M splits into two buckets. $100M is earmarked for coordination: convening global research institutions, standardising data collection protocols, and creating shared infrastructure for multi-modal biological datasets. The remaining $400M funds large-scale data generation, along with development of next-generation instruments for measuring, imaging, and engineering biology at cellular resolution.
All data generated will be open and freely available to the scientific community. Biohub is explicit that this is not a proprietary model-building effort — it is a data commons play, with the goal of reaching a scale no individual institution could fund alone.
Who Is Involved
The initiative has four major research co-ordinators at launch:
- Allen Institute
- Arc Institute
- Broad Institute
- Wellcome Sanger Institute
The Human Cell Atlas and Human Protein Atlas consortia are also participating. NVIDIA is named as the technology partner, providing accelerated compute infrastructure and domain-specific software for dataset processing and model training. Renaissance Philanthropy is joining to catalyse additional funding from external foundations.
Why This Is Hard
Predictive cellular models are a different order of difficulty from what AlphaFold accomplished with protein structure. AlphaFold operated on a well-defined problem with a specific output: 3D protein coordinates from an amino acid sequence. Cellular simulation requires modelling tens of thousands of gene products, their regulatory relationships, cell-type-specific behaviour, and disease-state perturbations — simultaneously and dynamically.
The data gap is the limiting factor. High-quality, multi-modal, single-cell measurements at the required scale do not yet exist in training-ready form. Biohub’s bet is that coordinated capital can create that foundation within five years, at which point the modelling problem becomes tractable.
Comparison to Prior AI-Biology Bets
Google DeepMind’s Isomorphic Labs ($600M commitment, launched 2022) targets drug molecule design. OpenAI’s science programme has invested in computational biology without publishing a dedicated funding figure. The Biohub initiative is structurally different: it is not building a commercial drug pipeline or a specific model, it is funding the open data infrastructure layer that any such model would train on.
At $500M for data generation alone, this is one of the largest non-commercial AI training data commitments on record.
No Outputs Timeline
Biohub did not publish a timeline for when predictive models would be available or evaluable. The five-year frame covers the data generation phase. Model development milestones were not disclosed.