Purchasing a book to burn it is a joke shared at a publishing world dinner party, yet has since been an effective engineering decision in the competition to create more successful language models.

Anthropic has legal filings that detail its internal process of “Project Panama,” which involved the printed books as a transitory container: buy in large quantities, peel off the spine, scan at industrial rate and then forward the remnants to recycling. The purpose of the project was sweeping as discussed in this aim- destroy all the books in the world- and the documentation indicates that the company did not want the project publicized. What lurks below the secret is a common conflict in the modern engineering world: the quickest pipeline is not normally the most beautiful, but it is the one that succeeds.
The logic is technical. Deep learning models trained on large text perform better when fed longer and more professionally edited text: text that bears a consistent framework, more tightly argued text, and has fewer of the hiccups that are the rule of web-scale chatter. In the filings of Anthropic, one of the co-founders expressed his opinion that books could be used to train models on how to write well, a reason that was also revealed in the inner communication of other companies. According to the court records, Meta referred to having access to enormous book collections as something vital to staying abreast. Books, that is, were turned into a value input material as useful as compute time.
The unorthodox method was the physical way. Anthropic employed a former Google executive, Tom Turvey, who worked on the Google Books scanning initiative, and scanned on a scale that was many times larger than most traditional digitization projects. One of the vendor proposals was that of turning 500,000 two million books in six months with a “hydraulic powered cutting machine” and production scanners and arranging to have them picked up by a recycling business. The destructive step is to achieve throughput: scanning pages that have been cut is both quicker and easier than scanning non-destructively captured pages, and also does not have to deal with handling constraints that slow down preservation-oriented workflows. The filings do not reflect the targeted collections of the rare ones, the documents refer to the bulk sourcing, including that of the big used-book sellers.
That branch of engineering was also a legal branch of engineering. A federal judge later found the process of training models off books could be considered fair use and defined training as “quintessentially transformative” and found, separately, that purchasing print copies and converting them into internal, searchable digital copies may be fair use as well. The most important boundary of the ruling was not the learning of the model, but the acquisition and storage of the inputs. The court made a clear differentiation between the act of digitizing legally acquired books and maintaining huge repositories of the digitized pirated copies.
The difference is important since the same filings represent a previous, less predetermined stage of data acquisition throughout the industry- one which was influenced by urgency and scale. Documents in the case of Anthropic indicated that co-founder Ben Mann would download books via LibGen, and then shared a link to Pirate Library Mirror with associates and wrote: just in time!!! Other instances explain such pressures. Meta workers talked in the company about the practice of torrenting book libraries, and internally one employee wrote that it “doesn’t feel right,” and that one of the issues could be to share pirated content in such a manner that the practice would be against the law. The courts have communicated that despite the possibility that model training can be fair use, liability can nevertheless be generated by downloading and retaining the pirated libraries.
Towards the beginning of 2026, this issue of liability has ceased to be in the background it has become a design constraint. An increasing group of lawsuits has been designed around the acquisition channel, namely, whether works originated in shadow libraries like LibGen, as opposed to the model-training act in its own right. Authors Alliance has followed 75 cases of copyright against AI since 2022, including book-specific cases that have proceeded through the fight over class certification, which can make the companies financially liable.
Viewed through the engineering prism, Project Panama verges on a weird scandal than a document-based portrait of an industry that had to learn to figure out how to work around the uncertainty in the law using hardware, logistics, and procurement. The action of purchase, cutting, scanning, and keeping within a company transforms the copyrighted text into what has proven to be treated by the courts differently in comparison to pirated caches. It also exposes a less pleasant tradeoff, namely that the physical book, as an object, is redundant once its words can be generated in bulk.
The future of the high-quality data could not be as midnight downloads but supply chains: contracts, warehouses, scanners, audit trails, and documentation that is sound enough to withstand discovery. The industrial process that nourishes a new model architecture is the most consequential innovation in that world.

