5 Sources
[1]
Harvard to share dataset of 1M public domain books for AI training
Harvard University has announced that it will be releasing about 1 million public domain books as a dataset available for the training of Artificial Intelligence (AI) models. This is set to be part of the Institutional Data Initiative (IDI), a program hosted within the Harvard Law School Library to
[2]
Google and Harvard drop 1 million books to train AI models
Harvard University, in collaboration with Google, will release a dataset of approximately one million public-domain books for use in training AI models, according to WIRED. This initiative, known as the Institutional Data Initiative, has secured funding from both Microsoft and OpenAI. The dataset
[3]
Harvard and Google to release 1 million public-domain books as AI training dataset
AI training data has a big price tag, one best-suited for deep-pocketed tech firms. This is why Harvard University plans to release a dataset that includes in the region of 1 million public-domain books, spanning genres, languages, and authors including Dickens, Dante, and Shakespeare, which are no
[4]
Harvard Makes 1 Million Books Available to Train AI Models
The dataset includes books that are in the public domain and no longer protected by copyright. Data is the new oil, as they say, and perhaps that makes Harvard University the new Exxon. The school announced Thursday the launch of a dataset containing nearly one million public domain books that can
[5]
Harvard adds copyright-free fuel to the AI fire.
With funding from Microsoft and OpenAI, the university's Institutional Data Initiative (IDI) is releasing a dataset for training AI that contains nearly one million public-domain books -- around five times larger than the controversial Books3 dataset. The aim is to "level the playing field" for
Share
Copy Link
Harvard University, in collaboration with Google, announces the release of a dataset containing approximately 1 million public domain books for AI model training, aiming to democratize access to high-quality training data for researchers and startups.

Harvard University has announced a groundbreaking initiative to release a dataset of approximately 1 million public domain books for training Artificial Intelligence (AI) models
1
. This project, part of the Institutional Data Initiative (IDI) hosted within the Harvard Law School Library, aims to expand and enhance the data resources available for AI training1
.The initiative is a collaborative effort involving Google, with the dataset comprising works from Google's extensive book-scanning project
2
. Notably, the project has secured funding from both Microsoft and OpenAI, highlighting the tech industry's interest in this resource3
.The dataset spans a wide array of genres, languages, and time periods, featuring classical texts from renowned authors like Charles Dickens, Shakespeare, and Dante, alongside more obscure works such as Czech math textbooks and Welsh pocket dictionaries
1
4
. This diversity aims to address the current limitations in AI training data, where various groups and perspectives are often underrepresented1
.Greg Leppert, IDI Executive Director, emphasized that the dataset is designed to "level the playing field" by providing access to a vast collection of high-quality training data for research labs and AI startups
3
. This move is particularly significant given the current landscape where AI training data often comes with a hefty price tag, favoring deep-pocketed tech firms3
.The use of public domain books circumvents the legal challenges faced by AI companies regarding copyright infringement. Recent lawsuits from major publishers against AI firms highlight the ongoing tensions in the industry over the use of copyrighted materials for AI training
2
. Harvard's initiative provides a legally safe pool of historical texts for responsible model training4
.Related Stories
While this dataset represents a significant step forward, questions remain about its sufficiency for comprehensive AI model training. The lack of contemporary references and updated language in these historical texts may necessitate additional data sources for AI companies seeking to create competitive and up-to-date models
2
4
.This project aligns with global efforts to ensure diverse representation in AI training data. For instance, Iceland has undertaken a national effort to ensure its language and culture are represented in AI models
1
. In India, plans are underway to launch an open-source forum, "IndiaAI Datasets Platform," by January 2025, aimed at hosting datasets for AI development1
.Summarized by
Navi
[5]
13 Jun 2025•Technology

26 Sept 2026•Technology

19 Nov 2024•Business and Economy

1
Technology

2
Policy and Regulation

3
Policy and Regulation
