DigestAI news desk

Cut through the AI noise.

Research

Study finds GPT-5.6-Sol offers competitive cancer survival estimates

Researchers introduced Survprompt, a framework that converts structured patient data into text prompts to test whether large language models can predict cancer survival without specialized training. The study evaluated frontier LLMs against conventional models, including random survival forests, using two multi-institutional pan-cancer cohorts: the public MSK-CHORD dataset and a new cohort from…

1 source primary source

Key points

  • Survprompt framework tests LLMs on survival prediction using MSK-CHORD and Providence St. Joseph Health data.
  • GPT-5.6-Sol achieved cMAE within 10% of specialized models for several cancer types.
  • LLMs showed inconsistent accuracy across institutions and poor discrimination between risk levels.

The results showed that GPT-5.6-Sol achieved censored mean absolute error (cMAE) within 10% of state-of-the-art specialized models for several cancer types. Notably, the LLM outperformed the specialized models for prostate cancer in the MSK-CHORD cohort. Feature ablations indicated that LLMs prioritized similar clinical variables as traditional survival models.

However, the study highlights significant limitations. LLMs displayed inconsistent accuracy across different cancer types and institutions. They also performed poorly in discriminating between high-risk and low-risk patients, resulting in lower concordance index scores. While zero-shot LLMs can generate surprisingly accurate prognostic estimates, their variable performance remains a barrier to clinical use.

Read the original at arXiv cs.CL · by Juan M Zambrano Chaves, Peniel Argaw, Risa Ueno, Carlo Bifulco, Kristina Young, Rom Leidner, Tristan Naumann, Hoifung Poon primary sourceOpen source ↗
Topics · follow one to build your own front page
Providence St. Joseph Health NetworkGPT-5.6-Sol

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories