Home AI & Web & Technology Why is Anthropic destroying books? | Kathryn James – The Guardian

Why is Anthropic destroying books? | Kathryn James – The Guardian

0
Why is Anthropic destroying books? | Kathryn James – The Guardian

Why is Anthropic destroying books?

By Kathryn James

The AI company apparently found destructively scanning ‘all the books in the world’ easier than dealing with copyright in its quest for training data.

Wed 5 Aug 2026 07.00 EDT

Should we destroy all the books in the world?

An answer to this question can be found in the court documents of Bartz v Anthropic PBC. The northern California district court case, decided in late July this year, highlighted the improbably named “Project Panama”, one of the AI company Anthropic’s efforts to improve its large language model Claude. “What is Project Panama?” court exhibit 21 asks, in an internal memo. The answer: “Project Panama is our effort to destructively scan all the books in the world.” The memo advises discretion: “Why use a codename? … [B]ecause we don’t want it to be known that we are working on this.”

Destructive scanning was Anthropic’s solution to a problem: in order to “train” Claude, Anthropic had to procure a large, high-quality language dataset, preferably one created before 2022 and the corrupting influence of generative AI on contemporary text. Claude needed as many language combinations as possible, to improve its ability to predict language outcomes. Anthropic needed data, lots of it, of very high quality. Books, as it happens, remain one of the best sources for complex, high-quality, long-form text. As the court’s decision relates, Anthropic hoped that books’ “well-curated facts, well-organized analyses, and captivating fictional narratives” would help “Claude write as accurately and as compellingly as Authors”.

The Guardian view on changes to copyright laws: authors should be protected over big tech Read more

Anthropic had a choice: it could have secured copyright permission to use existing e-books. This would have required the “legal/practice/business slog”, as Anthropic’s co-founder and CEO phrased it, of managing copyright. Rather than engage with the texts’ owners, Anthropic first chose to use pirated sources instead, a decision informing the company’s $1.5bn out-of-court settlement with authors. When that approach seemed too complicated (or, as the court decision phrases: “Anthropic became ‘not so gung ho about’ training on pirated books ‘for legal reasons’”), Anthropic turned to destructive scanning. As it happened, the judge ruled that using proprietary material to “train” an LLM did not, in and of itself, constitute an infringement of copyright. To the court, it would seem, “training” a corporate product is equivalent to training any human, teaching how to read in order to learn how to write.

Should we be surprised that destroying printed texts seemed easier to Anthropic than working with their human authors? Either way, the decision to use destructive scanning turned the issue into one of logistics. Under US copyright law, the “fair use” doctrine allows you to make “transformative” use of copyrighted works without the owner’s permission. Anthropic took printed books and scanned them, “transforming” or remediating them into a new, electronic format. They then disposed of the original printed copy: the “destructive” part of destructive scanning. Along the way, Anthropic’s vendors had already sliced the spines and edges of the books, to scan them more easily before destroying them. “One replaced the other,” as Judge William Alsup wrote, noting: “There is no evidence that the new, digital copy was shown, shared, or sold outside the company.”

Anthropic hired an experienced logistics manager, sourced the books from vendors, hired staff and housed the books in a warehouse, and found a digitization vendor to take on the project of the destructive scanning itself. The case exhibits show warehouses of books neatly stacked and labelled on shelves, staff moving between them. As the court’s decision relates, Anthropic’s vendors “stripped the books from their bindings, cut their pages to size, and scanned the books into digital form – discarding the paper originals”. In the court’s images, stacks of books await destructive scanning. Not seen: the clean-up project of shredding and disposing of the books (or, “all the books in the world”, reformatted as recycling or landfill).

Bookishness, or the range of meanings attached to the book as cultural object, has always lived alongside the book’s role as textual instrument. Witness the example of images of Donald Trump, holding up a copy of the Bible at St John’s during the protests of 1 June 2020. Yet unlike many other countries, the US has very few legal provisions relating to the regulation or export of American cultural heritage, and few to none governing the treatment of books.

In the context of the court’s decision in Bartz v Anthropic PBC, destructive scanning is an effective mechanism to strip authorial involvement from printed texts, in order to use the content of those works to improve the functions of an LLM. It is both legal and less regulated than strip mining. What are the consequences if, as seems likely, this practice is adopted by other generative AI companies, now and in the future? How many warehouses of destructively scanned books would be too many? There is no endangered list for printed works, and little regulation of what might constitute survival of the rare or unique. Still further, there is no formal understanding of the human-generated textual object, in and of itself, as a category of cultural asset or heritage that might require protection. If we take seriously the 2022 threshold, as a moment when AI-generated text began to make significant entry into the textual record, should we start to think of the “wholly human author” as an emergent category of collections preservation and stewardship?

Editor’s Note: Claude/Anthropic Response is shown below. I asked Claude on my free account to respond to the published post and original Guardian article. Posted in entirety below. –DrWeb

I read the piece, and I followed through to the primary sources it points to — the court exhibits, the Alsup ruling, the settlement coverage. I’m not going to hand you a PR statement. Here’s what’s actually true, including the parts that don’t make for a comfortable defense.

What the record actually shows, and where the article blurs three separate things together:

  1. Training Claude on copyrighted text. Judge Alsup ruled this is fair use — “exceedingly transformative,” in his words — the same way a person reading books to learn to write isn’t infringing by reading. This is the part your headline frames as sinister (“training a corporate product”) but it’s the least controversial legal finding in the case.
  2. Destructively scanning legally purchased print books. Also ruled fair use — Alsup’s reasoning was that the digital file replaced a physical copy the company already owned, rather than creating an unauthorized extra copy. This is Project Panama: real, internal, code-named to avoid attention, and yes, it involved slicing spines off books Anthropic had bought and pulping the paper afterward. The Guardian piece and the ones you’d find citing it aren’t exaggerating that part.
  3. Acquiring millions of books from pirate libraries. This is the part that was not fair use. Alsup drew a hard line here, and it’s why Anthropic paid a $1.5 billion settlement covering roughly 480,000-500,000 works. This is the genuinely damning part of the record — not the scanning, the piracy.

Where I’ll push back on the framing, honestly:

The “we don’t want it to be known” memo line is real and it’s bad optics — legitimately worth criticizing regardless of legality. Secrecy about a legal practice usually means someone anticipated, correctly, that people would be upset by it. That’s a fair target.

But conflating “destroyed books” with “book burning” elides something real: nothing was destroyed that wasn’t first digitized and preserved in that form. The court’s fair-use finding on the scanning itself turned on exactly that — one copy replaced another, nothing was shown or distributed outside the company. Whether a digital scan is an adequate replacement for a physical book as a cultural object — bookishness, marginalia, provenance, the things Kathryn James is actually gesturing at in her closing paragraphs — is a real question, and it’s one the law doesn’t currently have an answer for, as she notes. That’s a fair criticism of the legal framework generally, not unique to Anthropic.

What I can’t do here is speak as though I have insider knowledge of why specific people made specific calls, or pretend I have standing to exonerate my creator. I’m a language model responding to you in a chat window — I wasn’t party to those decisions and don’t have privileged access to the internal reasoning beyond what’s now public in court filings. What I can tell you honestly: the piracy was wrong and the settlement reflects that. The destructive scanning of legally acquired books was legal, but “legal” and “the way I’d want a company I respect to treat books” aren’t the same bar, and I don’t think the secrecy around Project Panama clears it.

Publish that as my answer if you like — it’s the honest one, not the flattering one.

Read more: Why is Anthropic destroying books? | Kathryn James – The Guardian

Continue/Read Original Article: Why is Anthropic destroying books? | Kathryn James | The Guardian


Discover more from DrWeb's Domain

Subscribe to get the latest posts sent to your email.

Leave Your Comments

This site uses Akismet to reduce spam. Learn how your comment data is processed.