DigestAI news desk

Cut through the AI noise.

Business & Funding3 min read

Oxford lets OpenAI train models on Bodleian library texts

The University of Oxford has permitted OpenAI to use digitised historical texts from the Bodleian library to train its AI models. While the partnership was announced in March 2025 primarily for digitisation purposes, internal documents obtained via a freedom of information request reveal that the scanned material was used to populate OpenAI's training set. By June 2025, 125,000 images from…

1 source

Key points

  • Oxford allowed OpenAI to use 125,000 digitised Bodleian images for model training by June 2025.
  • Internal minutes show staff concerns over reputational risk and environmental impact of the deal.
  • Oxford is the only UK partner in OpenAI's NextGenAI project, which includes US research libraries.

Staff and governance committee members raised concerns about reputational risks and the environmental impact of energy-intensive AI training. Oxford is the only UK member of OpenAI's NextGenAI project, which also includes US institutions like MIT and Caltech. An OpenAI spokesperson stated the collaboration helps preserve historical knowledge and ensures models reflect diverse cultures. The university maintains that the digitised material is out of copyright, modest in scale, and that the Bodleian retains rights to publish the scans openly online within months.

This move highlights a broader trend where AI developers seek fresh data from physical archives as scraped web content becomes saturated with AI-generated material. Unlike rival Anthropic, which has faced scrutiny for buying and pulping books for data, Oxford's deal keeps the physical collections intact. The university rejected claims that the training aspect was hidden, stating staff were aware the project would contribute training data.

Full story from The Guardian AI · by Ethan Penny and Dan MilmoOpen source ↗

Oxford lets OpenAI train its AI models on Bodleian library

The Guardian AI · 26 September 2026

Loading the full article…

This text was published by The Guardian AI and written by Ethan Penny and Dan Milmo. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗

Topics · follow one to build your own front page
OpenAIAnthropic404 MediaAmazonMarie EdgeworthDorothy Hodgkin

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Business & Funding

All →

Related stories