Posted on 07/30/2026 7:16:31 AM PDT by Red Badger
The ruling does change the fundamental nature of what the books were intended for: they were intended to be read by human beings, not used as training for machines to replace human thought and ingenuity.
==============================================================
AI companies are buying old books, scanning them, then destroying them. Anthropic has been buying the books to feed the machine with human knowledge and ideas written down before the AI era, so only books with publication dates before 2022 will do. Once they have the books, they rip off the spines, feed them through a scanner, and then on into the gaping maw of the AI LLMs that consume everything, including, eventually, themselves.
The revelations came out in a lawsuit brought by authors who said the destruction violated the Copyright Act. Federal judge for the Northern District of California William Alsup ruled that the bulk buy, scanning, and destruction is fair use under Section 107 of the Copyright Act. As far back as 2024, Anthropic, which has come under fire by the Trump administration over national security concerns, said "We don’t want it to be known that we are working on this." It's not a good look to be destroying books.
The suit reads that authors Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson were a few of the authors whose books were bought by Anthropic. For use in Claude, Anthropic "assembled these copies into a central library of its own, copied further various sets and subsets of those library copies to include in various 'data mixes,' and used these mixes to train various LLMs. Anthropic kept the library copies in place as a permanent, general-purpose resource even after deciding it would not use certain copies to train LLMs or would never use them again to do so." The authors had not agreed to this and these copies of their books were taken permanently out of circulation, never to be read by human eyes again.
Anthropic is one of these companies that has been buying physical books in bulk to rip apart, scan, then destroy, in service to its AI model Claude. Any books will do, even if that book is the very last copy of that book in the world. Legally, this is allowed under the first-sale doctrine, which permits a book-buyer to do whatever they want with that book object, without a need for the copyright owner's permission. It's an object, like a pair of shoes or a blender.
The ruling reads: "every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy. The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company."
In the suit, a judge found that because Anthropic was using the books in a way that fundamentally transformed them into something else, in this case turning them from books into raw data that will have no reference to their original self, and that this constituted fair use. Previous to the bulk buying of books, Anthropic pirated books or bought pirated books. They looked to former head of Google Books' scan project Tom Turvey to come in and get more data to feed the LLM. That data was not just information but also style and tone, proprietary authorial elements.
There's lots of companies that are willing to do the selling. ISBNdb claims to have the "world's largest book database," and says "the world's best AI training data is sitting on a shelf." They also promise not to reveal who is doing the bulk buying, knowing that no AI company wants to be the subject of a headline about how they're destroying millions of books at a go.
ISBNdb says that AI training on AI material results in data degradation. "Not all data degradation is intentional," says ISBNdb. "When AI systems train on text that was itself AI-generated, a documented phenomenon called model collapse occurs: subtle linguistic nuances vanish, systematic errors compound, and outputs converge on repetitive patterns. Each generation trained on synthetic data is slightly worse than the last. Print books from the pre-LLM era are structurally guaranteed to be free of this contamination. That alone is a significant advantage."
404 Media spoke to a bookseller who said that bulk orders from his store have resulted in the destruction of out-of-print books for which there are only single copies remaining. "I personally have mixed feelings about all of this," said the bookseller. "It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell. I’ve been well-suited for these sales with inventory from overseas and foreign language books. On the other hand, I don’t like the end-use, and I don’t like that uncommon books are being pulped."
The ruling reads: "This order grants summary judgment for Anthropic that the training use was a fair use. And, it grants that the print-to-digital format change was a fair use for a different reason. But it denies summary judgment for Anthropic that the pirated library copies must be treated as training copies." The ruling does change the fundamental nature, however, of what the books were intended for: they were intended to be read by human beings, not used as training for machines to replace human thought and ingenuity.
As it stands, Anthropic and other AI companies that are developing and furthering their large language models for use both as information and writing tools are permitted to take books out of circulation, feed them to the LLM, and destroy them, without concern for copyright law or, as it turns out, the preservation of the scope and breadth of human history, knowledge, experience, and creativity.
|
Click here: to donate by Credit Card Or here: to donate by PayPal Or by mail to: Free Republic, LLC - PO Box 9771 - Fresno, CA 93794 Thank you very much and God bless you. |
OOH! OOH! I want the signet ring!
I’ll find the actual link, get back to you.
I wonder how much it would cost to have a QR CODE ring made?.................
The source site overflows with fascinating old electronics books.
I bought my last tube-based device in perhaps 1973. They faded away pretty fast after that.
Bracelet or armlet might be better for that. I’m going to look into that idea of yours.
Old and historic books MUST be preserved physically—not just as “AI”.
This goes along with teaching young school children to READ—not just use cell phones!
Reading must be taught, preserved, practiced, and honored!!
I myself learned to read at age 4, since my parents provided me with a house full of elementary-level books and I am blessed with a high IQ. Today’s young children must have the opportunity to learn to read and do math.
And 16-19 year olds must have the opportunity to learn calculus! The “experts” who say that calculus is obsolete or unnecessary are full of baloney! Ask MIT—calculus and differential equations are still required at that university!
Who is claiming calculus is obsolete?
Thanks. One of these days I’ll get my old Telefunken Opus 7 back on line ...
Some nincompoop who regularly posts on “education” issues!
If they are purchasing the books, is it not their right to destroy their property? Or do we not own stuff we purchase?
If someone ones a book they can preserve it or destroy it...their call.
Don’t like it...you buy the book and preserve it.
If it’s rare it’s most likely out of copyright...so that’s a failed misdirection.
If they own the book they can chop it up all they want. If you don’t like it...go buy the book and you preserve it.
You said they don’t want to do the work that AI does. I wasn’t the idiot that said that.
I think after Dollartree they go to Ollie’s or Goodwill!.
It is what I would expect from a Clinton judge.
In some cases there are no other actual copies of the book available.
LOL. Descended from Ghengis?
I have a full set of encyclopedias bought in the 70s. Most never opened. I can’t get anyone to take them. No school or local library.
Not even Goodwill. I called them all.
I’m going to throw them away.
This seems to start here: The worry is that AI is displacing a lot of human work
The concept being discussed was the notion of transcribing the books into data files rather than using scanners, that taking away human jobs. I can assure you, software engineers do not want data transcription jobs where data is typed in from printed matter.
If you were thinking something else then "What we have here is failure to communicate".
Disclaimer: Opinions posted on Free Republic are those of the individual posters and do not necessarily represent the opinion of Free Republic or its management. All materials posted herein are protected by copyright law and the exemption for fair use of copyrighted works.