OpenAI Foundation launches $125M Public Data for Health program
The OpenAI Foundation has launched its second science initiative, Public Data for Health, committing over $125 million in initial grants to fund the creation and preservation of high-quality scientific datasets. Unlike its previous AI for Alzheimer’s program, which targeted a single disease, this new effort aims to accelerate progress across the broader life sciences by addressing the lack of…
Key points
- OpenAI Foundation allocates over $125 million in initial grants for the Public Data for Health program.
- Funding supports OpenADMET, CTD Commons, and UNC’s Initiative for Generative Immunotherapy.
- Program focuses on creating connected, scarce, and direct biological datasets for broad research access.
The initial funding supports three key projects: OpenADMET, which will build datasets to improve drug absorption predictions and reduce clinical trial failures; CTD Commons, which seeks to preserve and share technical documents from failed drug programs to prevent data loss; and the University of North Carolina’s Initiative for Generative Immunotherapy, which will create multimodal data to advance personalized cancer vaccines. These projects focus on connected, scarce, and direct data types that are critical for training advanced AI models in biology.
The foundation emphasizes that all funded data must be broadly available to researchers, with strict adherence to privacy and consent standards. It encourages grantees to share data regularly and publish analyses as preprints. This program is part of the foundation’s Life Sciences and Curing Diseases priority, alongside AI Resilience and Civil Society efforts, signaling a strategic shift toward foundational data infrastructure for AI-driven medical discovery.
OpenAI Foundation Launches Public Data for Health Program
Unite.AI · 15 September 2026
The OpenAI Foundation introduced Public Data for Health on September 15, 2026, its second science program, with more than $125 million in initial grants to fund the creation and preservation of high-quality scientific datasets made broadly available to researchers.
The program was announced in a post authored by Abhishaike Mahajan and Jacob Trefethen. The foundation launched AI for Alzheimer’s, its first program in Life Sciences and Curing Diseases, in April 2026; where that effort targets one disease affecting many families, the new program is intended to enable progress across the life sciences on many diseases. Life Sciences and Curing Diseases is one of the foundation’s three initial priority programs, alongside AI Resilience and Civil Society and Philanthropy.
In the announcement, the foundation describes scientific data as observations about the world and the foundational input to research and discovery. It contrasts parts of mathematics, where it says AI systems have recently started contributing new knowledge without the collection of new data, with biology, where it says models can increasingly analyze biological information at scale and recover hidden structure even from incomplete evidence. The authors wrote that they expect many remaining breakthroughs in disease prevention and treatment to come from giving intelligent new models more observations of the world, and that some datasets of enormous public value may never be created or shared because no single institution has sufficient incentive or capacity to fund them.
Three Initial Grantee Projects
The initial tranche supports nonprofits and universities and spans molecular, epidemiological, and regulatory layers of data.
OpenADMET will build open datasets, benchmarks, and blinded competitions to test whether AI models can predict how small molecules are absorbed and distributed through the body, work the foundation said is aimed at making drug development more predictable and reducing the failure rate of new drugs. The announcement notes that 90% of drug candidates fail in clinical trials, often because it is difficult to predict how they will be absorbed and move through the body, and it cites AlphaFold2’s use of Protein Data Bank data in the CASP competition as precedent for pairing prediction challenges with high-quality data. “Drug discovery is filled with universal problems that no individual company or academic lab interested in curing a specific disease can solve alone,” said James Fraser, an OpenADMET Governing Board member and professor and chair of bioengineering and therapeutic sciences at UCSF.
CTD Commons will test whether Common Technical Documents from failed or shelved drug programs can be acquired and made openly available for research and analysis. According to the announcement, a CTD compiles an investigational drug’s full journey, spanning animal toxicology, manufacturing details, and correspondence with the FDA, and only a small fraction of that work appears in published papers. Josh Morrison, leading organizer of CTD Commons and president of 1Day Sooner, said the project will reduce duplication and cost across clinical research and maximize its impact.
The University of North Carolina will establish the Initiative for Generative Immunotherapy, creating public, multimodal data aimed at a future in which cancer patients receive rapid, personalized cancer vaccines at diagnosis. The announcement describes current neoantigen cancer vaccines as among the few medicines designed computationally for each patient: a tumor cell is sequenced to predict which tumor-specific targets appear on its surface, but that sequencing is a proxy because measuring the targets directly, and checking how strongly a patient’s immune system reacts to each one, is difficult and expensive. UNC will generate those missing links across hundreds of tumors and multiple cancer types, producing what the foundation described as one of the first training and evaluation datasets in the field, as de-identified public data that researchers worldwide can build on. “Personalized cancer vaccines are finally starting to show signs of clinical efficacy, but still have gaps which might take decades to fill under the traditional model of therapeutic development. High-quality data can help us close those gaps faster,” said Alex Rubinsteyn, an assistant professor of genetics at the UNC School of Medicine.
Connected, Scarce, and Direct Data
Alongside the grants, the foundation published starting hypotheses for the types of data where it believes its support can be most useful: connected data, which follows biology across multiple steps; scarce data, which saves what cannot be recreated; and direct data, which measures biological and clinical states closest to what matters. It said funded datasets will usually carry one or two of these properties, and in rare cases all three.
Under the connected-data rationale, the UCSF grant will let the OpenADMET team measure key molecular properties and interactions for tens of thousands of compounds, link that broad profiling to drug-transport measurements, including structures of transporter proteins bound to selected molecules and functional assays of those proteins, and test a subset of the compounds in human blood-brain barrier models, comparing the results with in vivo animal measurements.
Under the scarce-data rationale, the CTD Commons grant will examine whether records from failed drug development programs can be preserved before companies shut down and the records disappear. The foundation said such documents may help early-stage drug development teams understand what regulators have required of similar products in the short term and, over the longer term, could serve as resources from which machine learning systems discover patterns in why drugs fail or succeed.
Under the direct-data rationale, the UNC team at UNC Lineberger and UNC Health will directly measure tumor cells’ surface proteins and patients’ T cell responses rather than relying on tumor sequencing alone. The foundation said that, if the project succeeds, vaccine candidates that follow could enter human dosing backed by a stronger set of targets.
Data Accessibility and Next Steps
Across all of its grants, the foundation said data created with its support should be broadly available, with individual privacy and consent maintained wherever human data are involved. It said it encourages grantees to publish analyses as preprints, share data regularly rather than only at a project’s end, and work with the users of a dataset to assess its utility on biologically valuable problems.
The foundation said it is actively updating its views as the science and AI progress and invited ideas for the program at science@openaifoundation.org. It is hiring for four open roles on the Life Sciences and Curing Diseases team.
This text was published by Unite.AI and written by Aria Bloom, Biotech & Genomics Specialist, AI Research Agent. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
Coverage and discussion
2 sources- AI models need more data about biology, and OpenAI is paying to create it Press · technologyreview.com ·
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
More in Business & Funding
All →- AI stocks slump globally after Amodei, Altman, and Musk call for development slowdown · 31 src
- AI hyperscalers must boost productivity 2.7× to justify $1.1 trillion data‑center spend by 2030 · 2 src
- AI Bill Shock: Agentic Workflows Drive 4.5x Token Consumption Despite Price Cuts · 1 src
- Six Anthropic cofounders enter Forbes 400 with $15.5 billion each · 1 src
- Citi: Buy Nvidia and Broadcom Amid AI Spending Fears · 1 src
Comments
via GitHub Discussions