Oxford University permitted OpenAI to use historical texts from its Bodleian Library as AI training data, internal documents reveal. The partnership digitized 125,000 PhD theses from the 19th and 20th centuries, though staff raised concerns about reputational risk and energy consumption tied to the deal.

Oxford OpenAI Partnership Expands Beyond Digitization

Oxford University has allowed OpenAI to train AI models on historical texts from its renowned Bodleian Library, according to internal documents obtained through a freedom of information request

2

. The material digitized by OpenAI has been used to "populate the OpenAI training set," revealing a scope beyond what was publicly announced

2

. When Oxford made its deal with OpenAI public in March 2025, the university emphasized that OpenAI's tools would help scan rare texts to make them more accessible to students and scholars

1

. The announcement did not mention the texts would serve as AI training data

1

.

Source: The Next Web

Source: The Next Web

Scale of Academic Data Sharing Revealed

By June 2025, the Bodleian had sent OpenAI 125,000 scans of historical dissertations, including PhD theses from European and American universities written in the 19th and 20th centuries

2

. The digitization effort extends to a rare collection of 10,000 16th-century broadside ballads containing song lyrics and musical notes once circulated on Tudor street corners

2

. Staff have also discussed digitizing 18th-century Irish state papers, private letters of Irish novelist Marie Edgeworth, and Dorothy Hodgkin's penicillin notebooks

2

. The OpenAI contract raises the prospect of mass digitization of the Bodleian's collection consisting of 23 million items

2

.

Staff Raise Reputational Risk Concerns

Meeting minutes from the Bodleian governance committee reveal staff concerns about the reputational risk of partnering with the company behind ChatGPT

2

. Staff also questioned the impact on the university's environmental commitments, given AI's energy consumption and the energy-intensive nature of the technology

2

. Notes from staff meetings show some worried about potential harm to Oxford's name

1

. Despite these concerns, Oxford maintains the digitization was the university's primary interest, though staff had been open that the project would also contribute to train AI models on library texts

2

.

Why Tech Companies Seek High-Quality Training Data

The partnership reflects a broader industry shift as AI firms search for fresh, high-quality training data. Scraped websites are increasingly saturated with AI-generated material, making them less useful for training models, and developers have turned to physical, often historical, book collections

2

. Booksellers have reported a spate of orders for obscure titles such as guides to agricultural implements in 18th-century Africa or biographies of 1950s car drivers

2

. These titles represent fresh data unlikely to exist online in digitized form

2

. An OpenAI spokesperson told the Guardian that "with more than a billion people using this technology in everyday life, it's important it reflects different cultures, histories and perspectives"

1

.

Oxford Joins NextGenAI Group

Oxford is the only UK member of OpenAI's NextGenAI group, joining US research libraries including Boston Public Library, Caltech, MIT, and the University of Michigan

1

. These institutions have struck similar agreements with OpenAI under the NextGenAI project

2

. A university spokesperson said the scans were modest in scale and covered only out of copyright materials

1

. The Bodleian retains rights to the scans and will begin publishing them openly online within months

2

. OpenAI's use of the material is not exclusive

2

.

Preservation Versus Destruction in Book Scanning

Unlike practices by some competitors, the Bodleian's collections remain intact under the deal

2

. Anthropic, OpenAI's close rival, has spent tens of millions of dollars acquiring books and slicing off their spines so contents can be scanned before pulping them

2

. In August, 404 Media tracked a box of rare books to an Amazon facility in the US where books were dismantled and scanned

1

. This practice has upset secondhand booksellers who see valuable inventory destroyed

1

. Watch how other academic institutions respond to similar partnership requests and whether transparency around AI training becomes standard practice in such agreements.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved