Oxford lets OpenAI train models on Bodleian library texts
The University of Oxford has permitted OpenAI to use digitised historical texts from the Bodleian library to train its AI models. While the partnership was announced in March 2025 primarily for digitisation purposes, internal documents obtained via a freedom of information request reveal that the scanned material was used to populate OpenAI's training set. By June 2025, 125,000 images from…
Key points
- Oxford allowed OpenAI to use 125,000 digitised Bodleian images for model training by June 2025.
- Internal minutes show staff concerns over reputational risk and environmental impact of the deal.
- Oxford is the only UK partner in OpenAI's NextGenAI project, which includes US research libraries.
Staff and governance committee members raised concerns about reputational risks and the environmental impact of energy-intensive AI training. Oxford is the only UK member of OpenAI's NextGenAI project, which also includes US institutions like MIT and Caltech. An OpenAI spokesperson stated the collaboration helps preserve historical knowledge and ensures models reflect diverse cultures. The university maintains that the digitised material is out of copyright, modest in scale, and that the Bodleian retains rights to publish the scans openly online within months.
This move highlights a broader trend where AI developers seek fresh data from physical archives as scraped web content becomes saturated with AI-generated material. Unlike rival Anthropic, which has faced scrutiny for buying and pulping books for data, Oxford's deal keeps the physical collections intact. The university rejected claims that the training aspect was hidden, stating staff were aware the project would contribute training data.
Oxford lets OpenAI train its AI models on Bodleian library
The Guardian AI · 26 September 2026
Loading the full article…
This text was published by The Guardian AI and written by Ethan Penny and Dan Milmo. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Business & Funding
All →- Anthropic’s founders seek voting control ahead of IPO · 5 src
- Meta's $14.3 billion Scale AI deal and Muse AI agent drive $234 billion market cap gain · 19 src
- Anthropic signs $11.6bn, seven-year cloud deal with Akamai · 14 src
- Google DeepMind alumni launch AI startups targeting alternatives to large language models · 1 src
- Confido raises $55M Series B led by Insight Partners · 1 src
Comments
via GitHub Discussions