Exploring a Books Data Commons for AI Training
Summary
Our work on copyright has long focused on supporting libraries and archives in the service of their missions to preserve and ensure access to culture. Among other things, that agenda calls attention to the ways in which copyright might impede libraries and archives who wish to make their collections available for research uses, including use for AI training in order to fulfill their public interest missions. on the availability and use of a dataset of books called “Books3” to train large language models (LLMs), a form of generative AI tool. We brought together practitioners on the front lines of building next-generation AI models, as well as legal and policy scholars with expertise in the copyright and licensing challenges surrounding digitized books. At the same time, as the paper highlights, there are already relevant examples of nonprofit and library-led efforts to provide responsible, fair access to books for many more people, not just the privileged few.