In early 2024, Anthropic launched a secret initiative code-named Project Panama—an audacious industrial pipeline to buy used books by the thousand, shear off their spines with hydraulic cutters, digitize the pages at speed, and recycle the remains, all to train its Claude AI model. Newly unsealed court filings and reporting by The Washington Post expose the scale of the operation, a $1.5 billion settlement with authors, and a federal judge’s nuanced fair-use ruling that together redraw the map for AI training data—and deliver a stark warning to enterprise IT leaders about data provenance, governance, and cost.
The Anatomy of a Book-to-AI Pipeline
Project Panama turned books into raw material. Anthropic’s workflow began with bulk purchases from used-book retailers like Better World Books and World of Books, court records indicate. Volumes arrived by the pallet and were fed to a hydraulic cutting machine that cleanly removed the spine, yielding loose pages for industrial sheet-feed scanners. OCR software converted the images to searchable text, which was then cleaned, deduplicated, tokenized, and fed into Anthopic’s training pipelines. The original paper was sent for recycling.
Internal planning documents, cited in the Washington Post, described the goal as an effort to “destructively scan all the books in the world.” A vendor proposal eyed a capacity of 500,000 to two million books in six months—an estimate, not a confirmed total, since many figures were redacted. The choice was aggressively efficient: non-destructive scanning is far slower and costlier, and storing millions of source books after digitization was avoided entirely.
For a company spending tens of millions of dollars, according to the unsealed material, the arithmetic was compelling. Used books can be cheaper than licensing digital editions outright, bulk procurement reduces administrative overhead, and a digital corpus is easier to filter, deduplicate, and push through machine-learning workflows. But efficiency, as the ensuing legal battle showed, is not the same as lawful or wise.
Why Physical Books? The Piracy Problem
Project Panama wasn’t a first move. According to Judge William Alsup’s June 2025 order in Bartz v. Anthropic, the company had previously downloaded at least five million book copies from the pirate site Library Genesis (LibGen) in June 2021 and at least two million more from Pirate Library Mirror (PiLiMi) in July 2022. Internal messages reveal co-founder Ben Mann personally ran the LibGen download over 11 days.
When Anthropic decided to move away from unauthorized collections, it hired Tom Turvey—who had led partnerships for Google’s book-scanning project—to help acquire “all the books in the world” while avoiding “legal/practice/business slog,” as the court order put it. Project Panama was the answer: buy legal physical copies, destroy them during digitization, and train on the resulting text without the taint of piracy.
The strategy mirrors a broader tension in AI. Books are prized training data because they offer coherent, long-form, editor-approved writing, internal documents noted—exactly the quality customers expect from a model like Claude. Yet licensing millions of individual titles is cumbersome and expensive. Physical acquisition, digitization, and destruction seemed a lawful shortcut.
The Fair-Use Ruling That Gave It Legal Cover
On June 2025, Judge Alsup ruled that using the books in question to train Claude was “exceedingly transformative” and thus a fair use under the specific facts of the case. Crucially, he drew a bright line: purchasing a physical book and scanning it for an internal digital library is legally permissible under fair use, while downloading and retaining pirate copies is not. The court said Anthropic lacked entitlement to keep the LibGen and PiLiMi collections, calling retention of that general-purpose library a separate, non-transformative act.
This distinction is the legal hinge. The ruling does not grant a blanket license to train on any text; it rewards lawful acquisition and internal, transformative use. It explicitly rejected the notion that piracy could be excused by an eventual AI-related purpose. The decision was bespoke, tied closely to facts—including the absence of proof that Claude generated infringing outputs and the destruction of physical source copies in the scanning pipeline. Future cases involving memorization, verbatim outputs, or different media could go another way.
$1.5 Billion Later: Lessons for AI Data Governance
The fair-use ruling did not end the matter. Anthropic later agreed to pay $1.5 billion to settle claims over the pirated files, without admitting wrongdoing. The settlement fund allocates compensation per covered title—reportedly several thousand dollars—and requires destruction of original files from the unauthorized sets.
For enterprise IT, the takeaway is blunt: “data first, permissions later” is a nine-figure mistake. A dataset without reliable provenance is a latent legal, financial, and reputational bomb, especially when it pollutes preprocessing systems, retrieval corpora, fine-tuning jobs, and model archives. Unwinding such contamination can become technically and legally impossible.
The settlement translates AI risk into a simple unit: liability per work. Organizations building internal AI tools must now treat training data with the same rigor they apply to software supply-chain security, open-source license compliance, and PII controls. Every searchable corporate archive—manuals, reports, customer documents—could become a model input. Records management, licensing, and retention policy are now foundational AI strategy elements.
What IT Leaders Should Do Now
For Windows and enterprise IT professionals, Project Panama is a wake-up call to action, not a spectator story.
- Audit your AI data pipeline immediately. Map every source of text used in training, fine-tuning, or RAG systems. Ensure you have the right to digitize, index, embed, and retain each document. Look especially at anything uploaded by employees to external AI tools.
- Insist on provenance tracking. As models move from research to production, maintain auditable chain-of-custody logs that show exactly where data came from, any transformations applied, and who approved its use. If a settlement demand arrives, you’ll need to demonstrate lawful acquisition.
- Review internal digitization practices. If your organization scans printed materials—manuals, periodicals, external reports—confirm that the activity falls within fair use or a license. The Anthropic ruling doesn’t automatically protect every in-house scanning project.
- Update your vendor due diligence. Ask any AI vendor or service provider hard questions about their training data sources, provenance, and copyright clearance. A model’s performance metrics don’t matter if its foundation is legally fragile.
- Prepare for regulatory attention. The $1.5 billion figure and the publicity around Project Panama will probably accelerate calls for transparency laws, data-provenance mandates, and tighter copyright enforcement. Start aligning your AI governance with emerging standards now.
The Cultural Cost: Scanned and Destroyed or Preserved?
Project Panama raises a deeper question beyond the courtroom: what is lost when a book becomes nothing more than fuel for an AI model? The destructive scanning model treats the physical artifact as disposable once its linguistic content has been extracted. For mass-market paperbacks with millions of surviving copies, that may seem unremarkable. But the risk escalates for out-of-print editions, scarce technical manuals, annotated copies, or culturally significant volumes that slip through bulk acquisitions without appraisal.
A scan captures words imperfectly; it does not capture typography, bindings, illustrations, marginalia, or physical provenance. The reporting does not establish that Anthropic knowingly destroyed irreplaceable books, but the workflow creates a structural risk: when a buyer acquires faster than it can evaluate rarity, history, or condition, preservation becomes an afterthought.
A responsible destructive-scanning program would need automated rarity checks against global catalogs, manual review for older or unusual titles, exclusion lists for special collections, and clear donation or resale channels. Without those safeguards, the drive for throughput can turn irreversible cultural loss into a routine operational detail.
Outlook: The Next Chapter in AI Data Sourcing
Anthropic’s Project Panama may become a template, a cautionary example, or both. The court ruling gives companies an incentive to buy physical books and scan them lawfully rather than rely on pirate libraries. But the $1.5 billion settlement and the public backlash over book destruction signal that licensing, governance, and cultural stewardship will only grow in importance.
Expect more AI developers to build in-house digitization pipelines, more publishers to offer granular AI-training licenses, and more litigation to test the boundaries of fair use with audiovisual works, software, and databases. For Windows and IT pros, the lesson is already clear: the most critical part of your next AI project won’t be the model architecture. It will be the data that feeds it—and the paper trail that proves you had every right to use it.