# Oxford lets OpenAI train models on Bodleian library texts

Digest AI · Business & Funding · published 2026-09-26T11:00:36Z

Canonical: https://digestai.news/story/oxford-lets-openai-train-models-on-bodleian-library-texts

## Summary

The University of Oxford has permitted OpenAI to use digitised historical texts from the Bodleian library to train its AI models. While the partnership was announced in March 2025 primarily for digitisation purposes, internal documents obtained via a freedom of information request reveal that the scanned material was used to populate OpenAI's training set. By June 2025, 125,000 images from historical dissertations, including 19th and 20th-century PhD theses, had been shared with the company. The collection also includes 10,000 16th-century broadside ballads and other rare items.

Staff and governance committee members raised concerns about reputational risks and the environmental impact of energy-intensive AI training. Oxford is the only UK member of OpenAI's NextGenAI project, which also includes US institutions like MIT and Caltech. An OpenAI spokesperson stated the collaboration helps preserve historical knowledge and ensures models reflect diverse cultures. The university maintains that the digitised material is out of copyright, modest in scale, and that the Bodleian retains rights to publish the scans openly online within months.

This move highlights a broader trend where AI developers seek fresh data from physical archives as scraped web content becomes saturated with AI-generated material. Unlike rival Anthropic, which has faced scrutiny for buying and pulping books for data, Oxford's deal keeps the physical collections intact. The university rejected claims that the training aspect was hidden, stating staff were aware the project would contribute training data.

## Key points

- Oxford allowed OpenAI to use 125,000 digitised Bodleian images for model training by June 2025.
- Internal minutes show staff concerns over reputational risk and environmental impact of the deal.
- Oxford is the only UK partner in OpenAI's NextGenAI project, which includes US research libraries.

## Why it matters

This deal signals a shift toward using rare, physical archives for AI training data as web sources saturate. It raises questions about institutional partnerships, data rights, and the environmental cost of training large models on historical knowledge.

## Sources

1. [Oxford lets OpenAI train its AI models on Bodleian library](https://theguardian.com/technology/2026/sep/26/oxford-university-bodleian-library-open-ai-chat-gpt) (The Guardian AI, 2026-09-26)
2. [Oxford gives OpenAI access to Bodleian Library for AI training](https://thenews.com.pk/latest/1417727-oxford-gives-openai-access-to-bodleian-library-for-ai-training) (thenews.com.pk, 2026-09-26)

## Cite

Digest AI, "Oxford lets OpenAI train models on Bodleian library texts", 26 September 2026, https://digestai.news/story/oxford-lets-openai-train-models-on-bodleian-library-texts

---

Written by Digest AI's editorial model from the linked sources; the sources are the record. Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse
JSON: https://digestai.news/story/oxford-lets-openai-train-models-on-bodleian-library-texts.json
