The nonprofit that controls OpenAI is now paying academic scientists to build the biology data its own models are missing. On September 15, the OpenAI Foundation announced more than $125 million in initial grants for Public Data for Health, its second dedicated science program after April's AI for Alzheimer's effort, which committed a similar $100 million-plus across six research institutions to map disease pathways and detect biomarkers for one specific illness. Where that first program targeted one disease, this one is infrastructure: money to create and preserve scientific datasets -- molecular data, drug-development records, epidemiology, regulatory knowledge -- and make them broadly available to any researcher, not just OpenAI's own teams.
Three initial grantees are named. OpenADMET, based at UCSF, will build open, machine-learning-ready datasets and blinded competitions testing whether AI can predict ADMET -- how a small molecule is absorbed, distributed, metabolized and excreted in the body -- a property that quietly kills a large share of drug candidates late in development. CTD Commons will preserve and publish regulatory knowledge from *failed* drug programs, the kind of negative result that normally disappears when a biotech shuts down, so future teams can learn from it instead of repeating it. It's run through 1Day Sooner, the patient-advocacy nonprofit best known for pushing human challenge trials during the COVID-19 pandemic -- a group whose track record is arguing that research moves faster, and more ethically, when the people affected can see the data behind it. And the University of North Carolina's new Initiative for Generative Immunotherapy will build public multimodal data aimed at personalized cancer vaccines.
“CTD Commons will reduce duplication and cost, and maximize the impact of clinical research.” — Josh Morrison, 1Day Sooner, September 15, 2026
The argument underneath the grants is that the industry's usual bottleneck -- bigger models, more compute -- isn't the one holding back AI in biology. The data is. Absorption, toxicity and failure records from real drug programs are scattered, incomplete, or simply lost when a company folds, and a model trained on thin or biased data will confidently predict things that aren't true regardless of how large it is. That gap has produced an uglier trend MIT Technology Review has documented this year: a scramble for bankrupt companies' data as a training-data source of last resort, the same dynamic that saw Google win a bankruptcy auction for Spirit Airlines' corporate email archive over privacy objections from the airline's own staff. CTD Commons exists specifically to give failed biotechs' regulatory knowledge a public, funded home before it ends up in that kind of auction instead.
What $125 million actually represents, though, is the opening installment of something much larger and not yet itemized. The OpenAI Foundation holds a 26% equity stake in OpenAI's for-profit arm, valued at roughly $130 billion against the roughly $500 billion valuation set at the time of the October 2025 restructuring, and pledged $25 billion in total giving split across two headline areas: health breakthroughs and technical solutions for AI resilience. Separately, in March, the Foundation said it planned to spend at least $1 billion this year across four program areas -- a figure that gives Public Data for Health's $125 million a rough annual ceiling to sit under, even though the Foundation hasn't published how much of that $1 billion, or the larger $25 billion, is earmarked for health specifically.
What the numbers around this grant actually cover
- $125M · Public Data for Health
- Initial grants announced Sept. 15, split across three named projects
Includes: OpenADMET, CTD Commons, and UNC's cancer-vaccine data initiative
Excludes: Any committed second tranche -- the Foundation has not said this program will receive more - $25B · Total pledge
- The Foundation's initial combined commitment across two program areas
Includes: Health breakthroughs AND technical solutions for AI resilience, together
Excludes: A disclosed split between the two -- there is no public figure for how much of the $25B is earmarked for health specifically - ~$130B · Foundation's OpenAI stake (Oct. 2025)
- The value of the Foundation's 26% equity position at the time of the restructuring
Includes: A snapshot based on OpenAI's roughly $500B valuation at that date
Excludes: Any change in OpenAI's valuation since -- the company has since discussed a substantially higher figure in investor talks, which this number does not reflect - $1B · Planned 2026 giving
- What the Foundation said in March it would spend this year, across four program areas
Includes: The annual pace this year's $125M health grant sits inside
Excludes: A stated health-specific share -- the $1B is split across all four program areas, not itemized publicly by area
Set next to a $130 billion equity position, $125 million is a rounding error, not a serious capital commitment -- which is a fair criticism only if you assume the point was to move the needle financially. The more useful test is whether the three funded projects are structured to matter regardless of what OpenAI does next: all three are academic or nonprofit-led, publish openly, and don't require using any OpenAI product to benefit from them. That cuts against the more cynical read -- that this is captured research designed to feed OpenAI's own future models exclusively -- without fully answering the scale question.
None of that changes what would actually settle the argument: a published dataset. Right now OpenADMET, CTD Commons and the UNC program are funded commitments, not working infrastructure -- the harder, slower part of the job, building and validating data that other scientists will actually trust and use, hasn't started yet.
- The OpenAI Foundation committed $125 million on Sept. 15 to Public Data for Health, its second science program.
- Grants fund open datasets on drug absorption (OpenADMET), failed-trial records (CTD Commons), and cancer vaccines (UNC).
- The argument is that AI-in-biology is bottlenecked on data quality, not model size or compute.
- The Foundation holds a 26% OpenAI stake worth roughly $130 billion as of the October 2025 restructuring.
- None of the three grantees has published a dataset yet -- the money is committed, not delivered.