Court documents made public during the Anthropic copyright lawsuit revealed how the AI company purchased and digitized millions of books, exposing a wider industry shift as developers turn to high-quality human writing to train the next generation of language models.
When books became AI's most valuable resource
The first generation of artificial intelligence's large language models learned primarily from the internet. Billions of web pages, online forums, academic papers and digitized archives supplied the text that powered today's AI systems. That abundance is beginning to give way to a different challenge.
A 2024 analysis by research institute Epoch AI estimated that if current scaling trends continue, frontier language models could be trained on datasets approaching the available stock of public human-generated text between 2026 and 2032. Although new online content continues to appear every day, the supply of high-quality human writing suitable for training advanced models is finite, prompting developers to look beyond the web.
One indication of that shift emerged from an ongoing copyright lawsuit against AI company Anthropic. In August 2024, a group of authors sued the company in federal court in California, alleging that it had used copyrighted books without permission to train its Claude family of language models. During discovery, internal documents revealed that Anthropic had purchased millions of printed books and digitized them into a permanent library for AI training.
The lawsuit exposed more than one company's data practices. As competition intensifies for high-quality human writing, governments, libraries, publishers and courts are confronting new questions about copyright, stewardship and who should control the written knowledge that underpins the next generation of artificial intelligence.
Beyond the internet
The first generation of large language models was built on internet-scale datasets drawn from sources including Common Crawl, Wikipedia, books and academic repositories.
Yet quantity alone has never determined a dataset's usefulness. The 2024 Epoch AI report argues that carefully edited books and reference works generally provide richer training data than large volumes of repetitive or low-information web content. Continued progress, the researchers argue, depends less on collecting more text than on finding higher-quality text.
Books help fill that gap. Unlike many websites, they contain sustained arguments, editorial oversight and material that has never been published online. The web is also becoming a more complicated training source. As AI-generated text proliferates across websites, developers face growing concerns about training future models on synthetic rather than human-created content. Researchers have warned that repeatedly training models on AI-generated outputs risks degrading performance over successive generations, strengthening demand for verified human-authored material.
A new economy of knowledge
Court filings show Anthropic's approach to acquiring books changed as it sought larger supplies of high-quality training data. According to internal planning documents, executives concluded that purchasing printed books and digitizing them would allow Anthropic to build a permanent research library while reducing its reliance on disputed digital collections.
Those records formed the basis of what The Washington Post later reconstructed as “Operation Panama,” an internal initiative to purchase millions of printed books, remove their bindings so they could be processed by high-speed scanners and convert them into a permanent searchable digital library.
Operation Panama marked a shift in strategy rather than the beginning of Anthropic's efforts to acquire books. Having previously amassed large collections from online repositories, the company turned to lawfully purchased print books to build a permanent research library, even as the legality of its earlier collections remained under scrutiny.
That distinction became central to the court's analysis. In June 2025, U.S. District Judge William Alsup held that Anthropic's digitization of lawfully purchased print books for internal AI training constituted fair use because it was sufficiently transformative. He separately ruled that the company's acquisition and retention of pirated digital books raised distinct copyright questions that would proceed to trial. Rather than settle the broader copyright debate, the ruling underscored that how AI companies acquire training data may prove as important as how they use it.
Who should steward the world's books?
The competition for books is also reviving a much older debate over who should control humanity's written knowledge in the digital age.
For years, discussions surrounding artificial intelligence and copyright have focused primarily on whether companies should be allowed to train models on copyrighted works without permission. The U.S. Copyright Office's 2025 report on generative AI training concludes that there is no single legal answer. Whether AI training constitutes fair use, it argues, depends on factors including the nature of the works, how they are obtained, the purpose of the use and its effect on existing or potential licensing markets. Rather than endorsing a blanket rule, the report envisions a future in which litigation, licensing agreements and technological innovation together determine how books enter AI training datasets.
Researchers argue the challenge is ensuring AI training collections are transparent, well documented and responsibly managed. In June 2025, researchers from Harvard Library and the Institutional Data Initiative released Institutional Books 1.0, a collection of 983,004 public-domain books spanning more than 250 languages that were digitized through Harvard's participation in the Google Books project. Unlike many AI datasets, every volume is accompanied by records documenting its origin, copyright status, digitization and processing, allowing researchers to trace exactly how the collection was assembled.
The contrast with commercial AI development is striking. Private companies view books as a competitive resource capable of improving increasingly sophisticated language models. Libraries, by contrast, have spent centuries preserving collections so they remain accessible, traceable and publicly accountable. The researchers argue those traditions should not disappear simply because books now serve as AI training data. Instead, they propose that libraries and other knowledge institutions help build AI datasets using those same principles.
The debate is unlikely to end with the Anthropic litigation. As developers seek larger supplies of reliable human writing, questions once confined to librarians and publishers are becoming central to artificial intelligence.
