Patient data may become key AI asset in healthcare, OpenAI, PurpleLab, Harvard projects say
OpenAI has begun linking its ChatGPT for Healthcare product with Epic’s electronic health‑record system, allowing authorized providers to pull read‑only patient information—notes, medications, labs and prior visits—into AI‑assisted workflows. The company says the integration could cut the time clinicians spend assembling fragmented records and also ties in public sources such as PubMed and…
Key points
- OpenAI linked ChatGPT for Healthcare to Epic EHR, giving read‑only patient data access for clinicians.
- PurpleLab launched CARES to enable reproducible research with real‑world health data across multiple specialties.
- Harvard team created Aladynoulli, trained on 683,000 records, to predict risk for 348 diseases using EHR and genetics.
PurpleLab announced the Center for Advancing Real‑World Evidence Studies (CARES), a platform meant to help researchers, academic groups and patient communities work with the firm’s large real‑world data sets. According to a company statement, CARES emphasizes reproducible, transparent methods and aims to support studies across oncology, cardiology, mental health and rare diseases, potentially turning routine health information into evidence that traditional trials struggle to deliver quickly.
Researchers at Dana‑Farber Cancer Institute, Massachusetts General Hospital and Harvard Medical School unveiled a machine‑learning model called Aladynoulli that predicts risk for 348 diseases using electronic health records and genetic data from more than 683,000 patients. While the model can update risk scores as patients age, the authors caution it has only been tested on historical data and needs further validation before clinical use. The article stresses that the promise of such AI tools hinges on trustworthy, secure, and responsibly governed patient data.
How Patient Data Is Driving the Future of Healthcare
Unite.AI · 17 September 2026
Healthcare’s AI race is increasingly moving away from the question of whether machines can generate medical information and toward a more consequential one: what happens when AI can work with the patient data that clinicians and researchers already have?
Three recent developments point to where that shift may be heading.
First, OpenAI recently connected ChatGPT for Healthcare with authorized Epic electronic health record data, bringing patient-specific clinical context into an AI workflow and potentially reducing the time clinicians spend searching through fragmented medical records.
Next, PurpleLab, a healthcare data and analytics company, just launched CARES, an initiative to support research using its vast amount of real-world healthcare data.
And finally, researchers at Harvard Medical School have developed an AI model that uses existing health records and genetic information to estimate the risk of hundreds of diseases.
Taken together, they suggest what many in the healthcare space have speculated for quite some time: that patient data could be one of the healthcare industry’s most valuable assets for innovation in care.
Let’s take a further look at each of these recent patient data breakthroughs.
From EHRs to AI Assistants
On September 1, OpenAI announced an Epic integration for ChatGPT for Healthcare that allows authorized healthcare organizations to bring patient information into AI-assisted workflows.
The integration can help clinicians review information such as medical notes, medications, laboratory results and prior encounters without manually piecing together a patient’s history from different parts of the record.
The system is read-only, meaning the AI does not independently modify the medical record.
That distinction matters. The immediate opportunity is not replacing the EHR or allowing an AI system to make decisions on its own. It is reducing the time clinicians spend finding and synthesizing information.
OpenAI is also connecting ChatGPT for Healthcare to official public sources, including PubMed, ClinicalTrials.gov, DailyMed and CMS data.
The bigger idea is the combination of patient context and external medical evidence in the same workflow.
Making Real-World Data More Useful for Research
PurpleLab, a data and analytics company that focuses on, among other things, Real-World Evidence (RWE) and Real-World Data (RWD) in the healthcare space, recently launched the Center for Advancing Real-World Evidence Studies, or CARES, with a focus on helping researchers, academic institutions, evidence networks and patient communities use real-world data more effectively.
The significance goes beyond simply making more data available.
Healthcare researchers have long faced problems involving inconsistent definitions, fragmented datasets and difficulty reproducing previous studies.
CARES is designed around reproducible research, transparent methods, academic partnerships and patient-centered research.
PurpleLab says researchers can work with its real-world datasets to study treatment patterns, outcomes and healthcare costs across areas including oncology, cardiology, mental health and rare diseases.
If successful, initiatives like this could help turn routinely generated healthcare information into evidence that answers questions traditional clinical studies cannot always address as quickly.
“Real-world evidence is strongest when it’s rigorous, peer-reviewed, and shared openly. With PurpleLab CARES, we’re demonstrating what our data can do in the hands of researchers, academic partners and patient communities,” Scott DuVall, Senior Vice President, Real-World Evidence, at PurpleLab, said in a company statement.
Predicting Disease Before It Appears
At Harvard, research teams are turning patient data into prediction engines.
A machine-learning model developed by researchers at Dana-Farber Cancer Institute and Massachusetts General Hospital can estimate a patient’s risk of 348 diseases using routinely collected electronic health records (EHRs) alongside genetic risk information.
The model, called Aladynoulli, was trained and validated using data from more than 683,000 patient records across three biobanks.
Researchers say it can dynamically update risk assessments as a patient ages.
Its potential is significant. Instead of looking at medical history only as a record of what has already happened, AI could increasingly use it to identify patterns associated with what might happen next.
But this is where caution is essential. The model has so far been evaluated using historical patient data, and further research is needed before determining how it performs in real clinical settings.
The Real Challenge Is Trust
The future of healthcare AI may therefore depend less on how impressive an algorithm sounds and more on whether the data behind it is accurate, representative, secure, and responsibly governed.
Patient data can reveal patterns that are difficult for humans to identify across years of medical history. But those same records contain some of the most sensitive information a person generates.
As healthcare organizations connect EHRs, research databases, and AI systems, the central question will not simply be whether AI can extract more value from patient data.
It will be whether the healthcare system can do so without losing patient trust.
That balance between better information and responsible use could ultimately determine how far AI moves from the research lab into everyday healthcare.
This text was published by Unite.AI and written by Adlin Pertishya, Reporter, Espacio Media Incubator. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
More in Research
All →- Experiment: Graph RAG outperforms standard RAG on multi-hop questions, but frontier models lead · 1 src
- Jeff Dean explains Gemini's multimodal origins and coding-driven reasoning gains · 1 src
- Google Research Unveils R4T: Diffusion Retriever Cuts Query Fan‑Out Latency 12‑20× · 1 src
- Blindspot Benchmark Tests Long-Horizon Safety of Tool-Using LLM Agents · 1 src
- CLEAR framework improves medical LLM accuracy via cross-source evidence adjudication · 1 src
Comments
via GitHub Discussions