Mark Zuckerberg and Priscilla Chan's philanthropic venture Biohub committed $500 million in April to a five-year effort to generate the data needed to build predictive AI models of human cells. On Oct. 7, that effort nearly quadrupled: the U.S. Department of Energy, the National Institutes of Health, and a trio of AI companies -- Google DeepMind, Isomorphic Labs, and Meta -- joined the Virtual Biology Initiative with a combined $1.8 billion in funding, data, and computing power.
The goal, in Biohub Head of Science Alex Rives's words, is a virtual cell -- a computational model accurate enough that a scientist could run an experiment digitally, predicting how a given cell responds to a drug or intervention before anyone touches a pipette. Rives calls it "one of the most important challenges for the next era of science."
An accurate predictive model of biology could dramatically accelerate scientific discovery by enabling scientists to perform experiments digitally.
The new money breaks down distinctly. The Department of Energy is putting in more than $500 million over five years, aimed at exascale computing, X-ray and neutron scattering, cryo-electron microscopy, and autonomous laboratories -- physical measurement infrastructure the private sector doesn't own. The NIH is contributing by coordinating datasets and repositories built from more than $500 million in prior federal biomedical research funding, rather than writing a new check. Google DeepMind, Isomorphic Labs and Meta are putting in $300 million in fresh corporate capital.
The NIH's piece is narrower than writing a check: much of the agency's existing genomic, imaging, and clinical-trial data was collected without a common format, so Kleinstreuer's agency is tasked with standardizing it under the same "Genesis Mission" banner as DOE's new measurement investment, so a model can actually train across datasets that weren't built to talk to each other.
Four funders, four different kinds of money
| Biohub committed April 2026 | Tech companies Google DeepMind, Isomorphic Labs, Meta | Dept. of Energy 5-year commitment | NIH prior investment | |
|---|---|---|---|---|
| Amount | $500M | $300M combined | $500M+ | $500M+ |
| What it is | New philanthropic capital | New corporate capital | New public funding (labs, compute) | Existing datasets/repositories, re-committed |
| Data access terms | Sets the open-data framework | Temporary embargo before public release | No embargo -- public immediately | No embargo -- public immediately |
The distinction matters for who gets the data first. Commercial funders receive an embargo period to work with the resulting datasets before public release, though neither Biohub nor any of the three companies has disclosed how long that embargo period runs. Data generated with government funding carries no such restriction and becomes public immediately.
- Generate multi-modal cell data using DOE compute, imaging technology, and Biohub-funded instruments
- Get an undisclosed embargo window to work with the data before anyone else
- Receives the same data once the embargo lapses, at no cost
Why these specific partners: Isomorphic Labs president Max Jaderberg said "generating the data to solve predictive systems biology requires scaling past the limits of what any single organization can produce today" -- an acknowledgment that even Alphabet-scale resources weren't enough alone. Google DeepMind vice president Pushmeet Kohli called the virtual-cell effort "one of the great collective scientific challenges." On the funding side, DOE Under Secretary for Science Darío Gil and NIH Deputy Director Nicole Kleinstreuer both tied their agencies' contributions to the government's existing "Genesis Mission" push to apply AI to public research. April's original commitment also drew the Allen Institute, Arc Institute, Broad Institute, and Wellcome Sanger Institute as research partners, with Nvidia signed on as a computing technology partner.
That reaction wasn't limited to the companies providing new money. When Biohub made its original commitment in April, Whitehead Institute/MIT biologist Jonathan Weissman called it "the kind of coordinated, openly shared infrastructure that can genuinely change what's possible in biology," and Human Protein Atlas co-director Emma Lundberg said a "global coordinated data foundation for modern AI-powered biology is exactly what we need to break siloes." The original $500 million itself was split specifically: $400 million for new measurement technologies like cryo-electron tomography and advanced microscopy, and $100 million to help coordinate outside research labs rather than fund Biohub's own work. The October expansion is an attempt to make that infrastructure large enough to matter before today's AI-biology efforts each lock in their own, smaller, proprietary datasets instead.
The scale problem is real, not a marketing line. Biohub says building an accurate virtual cell requires data spanning billions to trillions of individual cells and cellular states across different conditions -- orders of magnitude beyond the hundreds of millions of cells existing public atlases like the Human Cell Atlas and Human Protein Atlas currently cover. That gap is what the new money is explicitly meant to close, not a model-training budget.
- A head start on biological training data before anyone else gets it, in exchange for funding its generation.
- Eventually receive the same data for free once the undisclosed embargo lapses -- a resource no single lab could generate alone.
- The entire pitch is that a working virtual cell could compress years of lab experiments into simulations -- but Biohub's own scientists say the data needed to build one doesn't exist yet, at any funding level.
- A $1.8 billion, multi-institution open-data commons could become the default substrate other labs' models are trained and judged against, the way large open text corpora did for language models.
None of this produces a usable model soon. Biohub expects its first new datasets roughly a year out, and has said accurate predictive models remain a five-year target even with the expanded funding -- the same five-year horizon the project set for itself in April with a third of the current money. (A funding announcement measures intent and capacity, not progress -- the actual test of this initiative is whether the data-generation pace in a year matches today's numbers, not whether another funder joins before then.) The money to attempt a virtual cell now exists at a scale no single lab could match -- and the cell itself still doesn't.
- Biohub, the Dept. of Energy, the NIH and three tech companies committed a combined $1.8 billion total.
- Google DeepMind, Isomorphic Labs and Meta are jointly putting in $300 million in new money.
- The Department of Energy is adding more than $500 million over five years in labs and computing.
- Goal: a "virtual cell" model accurate enough to simulate drug responses digitally before lab testing.
- Caveat: commercial funders get a temporary head start on the data; no embargo length has been disclosed.