2 Sources
[1]
Oxford let OpenAI train AI models on Bodleian Library texts
Internal minutes show staff worried about reputational risk, while Oxford says the material is out of copyright. Oxford has let OpenAI use old texts from its Bodleian Library to train its AI models, according to internal papers seen by the Guardian. Ethan Penny and Dan Milmo broke the story on
[2]
Oxford lets OpenAI train its AI models on Bodleian library
University staff voice concerns over reputational risk of partnering with company behind ChatGPT The University of Oxford has allowed the company behind ChatGPT to train its AI models on historical texts from its Bodleian library, as tech companies scour academic institutions for fresh data. The
Share
Copy Link
Oxford University permitted OpenAI to use historical texts from its Bodleian Library as AI training data, internal documents reveal. The partnership digitized 125,000 PhD theses from the 19th and 20th centuries, though staff raised concerns about reputational risk and energy consumption tied to the deal.
Oxford University has allowed OpenAI to train AI models on historical texts from its renowned Bodleian Library, according to internal documents obtained through a freedom of information request
2
. The material digitized by OpenAI has been used to "populate the OpenAI training set," revealing a scope beyond what was publicly announced2
. When Oxford made its deal with OpenAI public in March 2025, the university emphasized that OpenAI's tools would help scan rare texts to make them more accessible to students and scholars1
. The announcement did not mention the texts would serve as AI training data1
.
Source: The Next Web
By June 2025, the Bodleian had sent OpenAI 125,000 scans of historical dissertations, including PhD theses from European and American universities written in the 19th and 20th centuries
2
. The digitization effort extends to a rare collection of 10,000 16th-century broadside ballads containing song lyrics and musical notes once circulated on Tudor street corners2
. Staff have also discussed digitizing 18th-century Irish state papers, private letters of Irish novelist Marie Edgeworth, and Dorothy Hodgkin's penicillin notebooks2
. The OpenAI contract raises the prospect of mass digitization of the Bodleian's collection consisting of 23 million items2
.Meeting minutes from the Bodleian governance committee reveal staff concerns about the reputational risk of partnering with the company behind ChatGPT
2
. Staff also questioned the impact on the university's environmental commitments, given AI's energy consumption and the energy-intensive nature of the technology2
. Notes from staff meetings show some worried about potential harm to Oxford's name1
. Despite these concerns, Oxford maintains the digitization was the university's primary interest, though staff had been open that the project would also contribute to train AI models on library texts2
.The partnership reflects a broader industry shift as AI firms search for fresh, high-quality training data. Scraped websites are increasingly saturated with AI-generated material, making them less useful for training models, and developers have turned to physical, often historical, book collections
2
. Booksellers have reported a spate of orders for obscure titles such as guides to agricultural implements in 18th-century Africa or biographies of 1950s car drivers2
. These titles represent fresh data unlikely to exist online in digitized form2
. An OpenAI spokesperson told the Guardian that "with more than a billion people using this technology in everyday life, it's important it reflects different cultures, histories and perspectives"1
.Related Stories
Oxford is the only UK member of OpenAI's NextGenAI group, joining US research libraries including Boston Public Library, Caltech, MIT, and the University of Michigan
1
. These institutions have struck similar agreements with OpenAI under the NextGenAI project2
. A university spokesperson said the scans were modest in scale and covered only out of copyright materials1
. The Bodleian retains rights to the scans and will begin publishing them openly online within months2
. OpenAI's use of the material is not exclusive2
.Unlike practices by some competitors, the Bodleian's collections remain intact under the deal
2
. Anthropic, OpenAI's close rival, has spent tens of millions of dollars acquiring books and slicing off their spines so contents can be scanned before pulping them2
. In August, 404 Media tracked a box of rare books to an Amazon facility in the US where books were dismantled and scanned1
. This practice has upset secondhand booksellers who see valuable inventory destroyed1
. Watch how other academic institutions respond to similar partnership requests and whether transparency around AI training becomes standard practice in such agreements.Summarized by
Navi
[1]
[2]
13 Jun 2025•Technology

23 Jul 2026•Policy and Regulation
02 Apr 2025•Technology

1
Policy and Regulation

2
Technology

3
Technology
