
In recent months there has been a growing alarm call in the book world, leading to such visceral headlines as “the Library of Alexandria burns again”, “AI companies are pulping our humanity”, and a whole host of monster movie descriptors – Ingesting! Destroying! Gobbling!
The word from second-hand book sellers¹ around the world is that they’re receiving bulk orders from mysterious companies for large but random selections of used books. By large we mean anything from the hundreds to the thousands. The belief is that the secretive entities are destructively scanning the books to train AI LLMs², then pulping them.
To give some background into how this has come about, we have to look at the part of copyright law usually associated with Weird Al Yankovic parodies and such like.
TRANSFORMATION!
In United States copyright law there is something called ‘transformative use’ which means you can use a copyrighted work as the basis for your own, providing you impart new meaning or value onto it. It covers parody but thanks to AI, increasingly has to deal with much more complicated scenarios.
One of these scenarios concerns Google Books³. In the early 2000s, Google started scanning public domain books to create a searchable online library. Then, in partnership with ‘real’ libraries, it increased its scope to include copyrighted works, resulting in the 2015 court case Authors Guild v. Google. Google was charged with infringement of copyright by providing access to books still under copyright without the authors’ permission. Google won the case as the process of scanning, as well as the search and snippet functions, were deemed to be transformative. Google was providing information about the book but not allowing the copyrighted book to be read in full.
INFRARED SCANNING!
Where Google had to be commended, was for its physical process of digitising the books. As anyone who has ever scanned pages from a book knows, it’s a laborious process of squishing the book between the scanner plate and lid and trying to find that happy place of capturing nice flat pages without damaging the book’s spine. Google developed its own patented method of scanning using infrared lasers, overhead cameras and human page turning which enabled it to see curved pages as flat and scan up to 6,000 pages per hour. Out-of-print books were borrowed from partnership libraries (in some cases, being shipped across the world to Google’s US facilities), scanned and returned. New in-print books were scanned using destructive methods (more on that method in a minute).
Do you remember our feature on Ai Weiwei and collecting in quantity in issue 118 (summer 2025)? We looked at Ai’s willingness to destroy Chinese antiquities because there was an abundance of them. Ai said, “it’s not like I destroyed a Rembrandt, it’s not like there is a limited number.”⁴
What would come to follow the precedent set in the Google case was a number of AI companies looking to obtain the breadth and quantity of content, but without the same intention to share what they’d assimilated. The secrecy of their operations afforded them the opportunity of using less ethical methods of achieving their end goals.
PIRACY!
The court case at the centre of the current controversy is Bartz v. Anthropic. It was a US class action taken out by authors against AI company Anthropic⁵. Anthropic had downloaded approximately 7 million pirated books from LibGen⁶ to train its LLM, with pre-2022 content being particularly valuable because it is guaranteed to be free of AI-written content. Another court case, Kadrey et al. v. Meta Platforms, dealt with the same mass piracy. While you would imagine such huge corporations would have the money and means not to resort to piracy, court documents revealed that for Meta, piracy was preferable than “a shit ton of work”⁷ having to license books from individual authors.
Anthropic agreed to pay $1.5 billion to settle the piracy aspect of the lawsuit but, crucially, the judge ruled that using books for the act of AI training was fair use providing the books were obtained legally.
BULK PURCHASING!
As the court order would reveal, Anthropic would become “not so gung ho about” training on pirated books “for legal reasons”. So Anthropic hired the former head of partnerships from the Google project, Tom Turvey, and gave him the job of obtaining “all the books in the world”. Turvey initially made contact with publishers to enquire about licensing, but this fizzled out and his team moved onto the idea of bulk purchasing under the codename of “Project Panama”. Anthropic secretly (hence the codename) spent millions of dollars to purchase millions of books, beginning a process of stripping the spine, cutting pages to size, scanning them and destroying the paper originals.
Investigations by The Atlantic into the more recent cases of bulk purchasing have not resulted in a concrete answer as to who is behind the purchases, what the books are being used for or, perhaps most importantly, what is happening to the books afterwards, but the view is that it is almost certainly the work of AI firms. What is also unknown is how rare the purchased words were – is the information contained in some books lost forever? With the purchasing being made on such a massive scale, we must presume that no checks are being made to ensure that any last remaining copies aren’t being removed from circulation.
BOOK BURNINGS!
The intent of this mass destruction of books is the opposite of the censorship-driven book burnings of the past, but the outcome is the same – content (potentially) being lost forever. Maybe, if these books have been languishing on a bookseller’s shelf for years, they aren’t particularly valuable to us in terms of their content, but are we comfortable with an unscrupulous AI firm taking that decision out of our hands?
It is a difficult ethical question for booksellers, whether to hold onto stock that may never sell, or sell it and know that the dispatch address could be the ferry to the underworld.
Next time you buy a second-hand book, don’t feel guilty about the strain of the extra weight on your bookshelf, or worry that it’s just going to collect dust – you could be saving it from the Grim Reaper.
Notes:
- A Dutch bookseller interviewed by Fortune received a request from a company called 2077AI to send 3,001 mostly academic publications to an address in China.
- Artificial Intelligence Large Language Models. A system that has been fed vast amounts of text to teach it to understand/generate language.
- Formerly known as Google Print, then Google Book Search. It allows users to search the text of millions of books, showing snippets in its results (much like a standard Google search). Only the snippets are viewable for copyrighted works, but the full book text can be viewed on public domain works or where permission has been granted.
- Tiffany Wai-Ying Beres, “A Battlefield of Judgements: Ai Weiwei as Collector,” Orientations 46, no. 7 (2015), pp. 87–92
- Anthropic is a firm backed by Amazon and Google’s parent company, Alphabet. It is the company behind the Claude AI assistant.
- Library Genesis. A Napster-style file-sharing library, providing free and unauthorised access to books, probably the largest online library in the world.
- Kadrey et al. v. Meta Platforms, Document 417-6
Additional sources/further reading:
The Atlantic: Someone Is Mysteriously Snapping Up Used Books Around The World
Authors Alliance: Looking Back at Google Books Eight Years Later
NPR: The Secret of Google’s Book Scanning Machine Revealed
