(annas-archive.pk)
https://en.wikipedia.org/wiki/Google_Books
https://arstechnica.com/tech-policy/2015/10/appeals-court-ru...
https://arstechnica.com/ai/2025/06/anthropic-destroyed-milli...
I thetefore feel the same way about Google Books that how I felt when I learned that What.cd went down: that I don't gain or lose anything anyway because I never had access to begin with, and that by not making it 100% publicly accessible you're asking for the data to one day disappear forever.
Big companies will read up the books and make their AI recite them from memory, but Archive.org was sued for renting one book on an exclusive basis (unless one would return, another wouldn't be able to rent)
No, this is what they were doing before, but they explicitly started lending out "unlimited" copies, which is why they got sued.
> “At bottom, [the Internet Archive’s] fair use defense rests on the notion that lawfully acquiring a copyrighted print book entitles the recipient to make an unauthorized copy and distribute it in place of the print book, so long as it does not simultaneously lend the print book,” Judge John G. Koeltl of the U.S. District Court in Manhattan wrote. “But no case or legal principle supports that notion. Every authority points the other direction.” [0]
[0]: https://www.insidehighered.com/news/tech-innovation/teaching...
> The crux of IA's first factor argument is that an organization has the right under fair use to make whatever copies of its print books are necessary to facilitate digital lending of that book, so long as only one patron at a time can borrow the book for each copy that has been bought and paid for. See Oral Arg. Tr. 31:10-15. But there is no such right, which risks eviscerating the rights of authors and publishers to profit from the creation and dissemination of derivatives of their protected works. See 17 U.S.C. §§ 106(1), (2). IA's wholesale copying and unauthorized lending of digital copies of the Publishers' print books does not transform the use of the books, and IA profits from exploiting the copyrighted material without paying the customary price. The first fair use factor strongly favors the Publishers.
> In this case, there is a "thriving ebook licensing market for libraries" in which the Publishers earn a fee whenever a library obtains one of their licensed ebooks from an aggregator like OverDrive. Pls.' 56.1 ¶¶ 577-578. This market generates at least tens of millions of dollars a year for the Publishers. Id. ¶¶ 170, 172. And IA supplants the Publishers' place in this market. IA offers users complete ebook editions of the Works in Suit without IA's having paid the Publishers a fee to license those ebooks, and it gives libraries an alternative to buying ebook licenses from the Publishers. Indeed, IA pitches the Open Libraries project to libraries in part as a way to help libraries avoid paying for licenses. See Pls.' 56.1 ¶ 383 (presentation IA gave to libraries asserting that pairing with IA means that "You Don't Have to Buy It Again!"); id. ¶ 382 (different presentation promising that the Open Libraries project "ensures that a library will not have to buy the same content over and over, simply because of a change in format"). IA thus "brings to the marketplace a competing substitute" for library ebook editions of the Works in Suit, "usurp[ing] a market that properly belongs to the copyright-holder."
I don't think they were overcome. As far as I remember Google couldn't make the books available so they abandoned the project. They possess the scans (if they didn't delete them) but they won't be made public.
It's great they did this, but the Google that is today cannot be trusted with data of public value anymore.
Book scans, secreted away, are worthless to the public.
Info on where to send books not yet in their collection: https://help.archive.org/help/how-do-i-make-a-physical-donat...
Mobile apps to determine if they need a book: https://help.archive.org/help/donate-books-app-for-ios-and-a...
Web app: https://archive.org/want/?mode=donation_book
For example, I donated a copy of Systems Bible (out of print, hard to find imho) and paid for it to jump the digitization queue (https://archive.org/details/systemsbiblebegi0000gall/). The original book will remain stored as a physical backup. It's not fully publicly available of course due to copyright (it will eventually be made public by the Internet Archive once its copyright expires ~2084 and it enters the public domain), which is where shadow libraries|archives like Anna's Archive and Z-Library fill the gap.
If you have rare books you would like digitized, archived, and distributed, I am very interested in providing assistance.
Just taking one of those and "destroying them"(it is not destroyed, a digital copy with the ability of doing millions of copies is stored somewhere) is not problematic for Humanity.
By the way, I always search for second hand books. Most of the books there are garbage. Most people clean their shelves with the books they don't care about, but preserve the ones that are good. If they are young people that inherited a house and don't care about books, they pick and sell the good ones, giving away the bad books.
If you go to a recycling centre, the garbage to quality ratio is over 100 or more. That is, for every 100 books that are garbage there is one good quality book. It is very rare to find a jewel there.
somewhere were we can't access it. as the article states: “permanently locking human knowledge inside private corporate servers”
the issue isn't that the physical copy is gone, it's that they are preventing people from making digital copies that are actually accessible by destroying the physical copies.
> If you go to a recycling centre, the garbage to quality ratio is over 100 or more.
archivists keep everything, because we don't know right now what will be important 100 years from now.
By any metric imaginable, it's making the information more accessible, not less. First, it's taking a single copy of a 10k physical print and it's making it digital. Is it "locked"? Yes, by copyright laws, if you don't like that lobby to have them changed. But it's _closer_ to being widely available, not farther.
Plus having the info part of a LLM makes it immediately available to literally billions.
I happen to actually actively shop second hand bookstores, so I am potentially affected by this - as opposed to most people complaining because they don't like the idea. And I still absolutely support it.
Help me understand how. Not only are these LLMs expressly prohibited from specifically regurgitating copyright works if the users asks them to, but they habitually hallucinate or paraphrase things wrong.
If they won't regurgitate the copyrighted text verbatim, and are known to be confidently incorrect and hallucinatory, I'm struggling to see how these texts are "immediately available to literally billions".
And I ask this as someone who has had LLMs give me incorrect assertions about the contents of books.
For example, knowledge about how the book smells when you open it, something that reader do talk about enough I don't think this should be a strawman, is lost. But, that is about the experience of reading the book, not the knowledge of the book.
The exact text? Yeah, I think that is largely lost as well. This is a summary. And for rarer books, it will be a particularly bad summary. The basics of the book are being captured in a space that the right question that retrieve it, but worse than a sparksnote and any well read reader will tell you all the sorts of things a sparknotes already loses compared to reading the book directly.
But, that little bit of data is a bit more data than existed before, and future LLMs should get better at giving the information. So, in that sense, the knowledge is better being spread compared to copyright where the book stays in a warehouse until it is disposed of. If it was between this and a sparksnote of the book being made, the sparksnote is far better, but between this and the book simply being disposed of, then the LLM is better but far, far from great.
That's a lot of assumptions that goes into the judgment, which is probably why different people reach different conclusions. One person is imagine the alternate fate of this book being slowly rotting in a landfill, the other resting on a bookshelf where it is read at least once fully and then flipped through time to time, and neither are wrong.
> Plus having the info part of a LLM makes it immediately available to literally billions.
Isn't this a contradiction? I mean, maybe you can invoke a fair use policy if an LLM spits out some text from the scanned book, but then you are _not_ "making it available to literally billions".
History tells us that very few "permanent" situations are truly permanent.
Provided a set of information has value (which in this case it clearly does) then the overwhelming likelihood is that, eventually, through some method or other, the information will become public.
If you destroy the only copy of a physical artifact, the situation is as permanent as it can get.
Fetishizing books isn't going to help.
In fact, most of those "rare" books don't sell because nobody wants them. The AI companies are making them even more rare, so the booksellers should be thankful.
But where did you hear that they’re buying “all copies”? And to what end?
To take a crude example in a different medium - new york in the 90s - widely documented right? Yet, if you want to find high definition video of street life in a given burrough on a given day or year, you're faced with an enormously difficult task. There were some HD test videos done in Manhattan in the late 90s (which have been posted to Hackernews before), but there's no equivalent for the other burroughs. Your best best would be finding original negative out takes or location scouting footage from feature films, a very hard task. That's only 30 years ago. Outside of the focal points of the worlds attention - English language, rich countries, places in the news, contemporaneous sources for 'non notable' events (lifestyle, how people spoke dressed etc) is surprisingly poorly preserved.
Hopefully you can infer how this tracks to the written world and primary sources for language, technical manuals etc etc.
I honestly can't and I think you can't either or you would have used an example with books/printed media rather than film, an entirely different ballgame.
Similar textual examples would be any text containing actual language as it's spoken in a given place or time. Or any factual textbook detailing the buildings present in a given location. Or any text book detailing a now defunct construction process. Or any text book (generally small run) detailing a niche interest, now missing ecosystem or the state of a particular political situation at a given time. Essentially all textual primary sources for events which are not currently considered important - but which we have no way of estimating the future importance of. One can continue to create countless counterfactual examples in this vein. My overall point is we cannot know what may be useful or even essential in the future, and knowledge should be preserved under the assumption that it is likely to be.
Historians frequently refer to this paradox - how everyday aspects of life are frequently not explicitly documented, since they're so obvious to the communities or communities of expertise that observe and carry them out. So it's actually incredibly important to preserve what seems like ephemera.
Hell we couldn't have AI training at all if we lacked the corpus of existing written literature - but there was no way any author could have anticipated this future utility more than a couple of decades ago.
But buying all the copies is categorically different and I cannot imagine why they would attempt that, except perhaps to prevent their competition from getting a hold of the same text?
Of course they aren’t buying “all copies”, and that wouldn’t even be possible in most cases since such books are usually flea market/attic material and most copies aren’t for sale (or even catalogued) to begin with.
I’d be interested to learn who comes up with such lies though. Is it really just random people venting their frustration, or some kind of organized astroturfing operation?
Maybe the Bodleian Library at Oxford University should give all their rare books to AI companies to scan and destroy them since it is not a "big deal" anyway.
Except that when they did do a pilot with OpenAI to scan these rare books, [0] they did NOT destroy them. I wonder why?
[0] https://www.bodleian.ox.ac.uk/services/research-partnerships...
From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge.
Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available. As stated in my first paragraph, they are not the exclusive stewards of humanity despite them anointing themselves as such. Granted, a lot of these books might not be that useful, but still a relic of times pre-machine generated text, which makes them valuable if only for their archival value.
How tho?
Ever sunday at the flea market, I see thousands of books that are rotting, hoping for someone to buy them or at least take them home, so the seller doesn't have to pack them for the trip back. Just the other day, there was a whole bin of books in front of a shop, offering them for 50 cents a piece. They will be destroyed anyways.
Unless they are buying and destroying really old, rare books or important small-print books, it is not much damage. It is not like they will buy "all copies of all of the books", just one. And its just that the data in physical print most likely hasn't been used for training, so this can help you find more unmined quality data. Nobody is stealing your books, preventing you from buying more or destroying all copies of a single book.
And some of these books would rot out of circulation or be destroyed anyways. Some people throw away 80-100 year old books on the regular, as they might just be unimportant to them or the world in general. And once the last copy is thrown or rots, that book will die forever. This way, it will live forever instead, scanned and trained on, conjoined with the rest of our knowledge in a magic machine.
Liberians have to accept that much of their job is sending books to be burned. A lot of them try their best to get people to be interested in older books, but they have to make way for "newer" books instead.
The article is written in the same way as how dogs are getting murdered in the dog shelter, even if "everyone" agrees it is wrong, yet nobody adopts them.
The internet made this even worse by celebrating people who only have to say the right thing.
The major lesson I myself learned as I got older is that “you should do the right thing” is a rephrasing of “you should do what I want you to do”.
The internet has amplified voices who say things to signal rather than do things to change.
That's orthogonal to knowing what "the right thing" is.
> The example book of Old books of agriculture is probably not that important today
If I may be flippant, not to you but to the sentiment, skill issue.We're about to enter an era of climate instability that's going to cause wild fluctuations in the ability to grow food across the globe. Historical agriculture data AND data about confounds is crucial for figuring out what strains outside of our current mostly mono-strain agricultural supply chain could be cultivated.
And that's just one use case out of thousands; what if you want to understand and reconstruct technology adoption from that era?
What if... you just want to learn what your ancestor was doing at such and such time?
What if you want to find clever techniques for robot arms to work with food crops in space?
Or, heck just the alpha from a hedge fund point of view of finding old climate patterns and... :)
Your ability to make the most of knowledge is only limited by your imagination.
> Liberians have to accept that much of their job is sending books to be burned. A lot of them try their best to get people to be interested in older books, but they have to make way for "newer" books instead.
I don't understand what you're trying to say here. > The article is written in the same way as how dogs are getting murdered in the dog shelter, even if "everyone" agrees it is wrong, yet nobody adopts them.
https://en.wikipedia.org/wiki/No-kill_shelterre: saving books, at a personal level, I try to use the excuse of work to find, read, and do stuff with old books,
https://1517.substack.com/p/powder-and-stone-or-why-medieval
And yes, people still care. And people who care do things.
"Hey, guys, when the librarians get pissed about the destruction of books, it’s time to put those listening ears on.
Because we are very comfortable with the idea that books are tools that can be retired. What’s happening right now is not that...."
https://bsky.app/profile/annabookwriter.bsky.social/post/3mt...
I don't think many Liberians like the idea that they have to do it; it is just one of those things that has to be done, sadly.
It is pretty standard procedure.
But that's just it, it won't live on forever because, from the perspective of preservation, training is a lossy, noninvertible transformation. The LLM cannot legally produce the book verbatim, it will only spit out a regurgitation of the information, chopped and mingled into a broad information space.
Furthermore, these "magic machines" are not the property of the public. They are owned by a handful of corporations who want to charge you continuously for every token output by the machine. So, not only is the original text locked away forever behind company walls, you now need to pay for access to an approximation of the original contents which you can no longer even verify as being correct because the source is no longer accessible.
If you are cool with this, from a cost perspective you are cool with a deal whereby I trade you access to a definite resource for a one time fee of $N for, instead, a perpetual cost of $M to you every month/day/hour for access to an amalgam in which you cannot even determine what proportion of the resource you are actually getting. You're basically saying you're cool with me selling you some unknown portion of wine for a monthly subscription price instead of selling you a definitive amount of wine for a one time fee. lol.
But unreadable by humans, right?
Keep in mind we wouldn't even know this was happening were it not for investigative journalism.
Pretty feeble investigative journalism if they cannot give us the names of even 5 such books we are supposed to be outraged about.
> Correct. I work for a large used bookstore with an online component. We're getting slammed with orders for books like the proceedings of an obscure 1992 Dutch geology conference or $500 festschrifts about D-module applications we would have previously sold to some university library. We've never once had an order for anything anybody would actually want, and most of this shit has sat on our shelves for years, if not decades. It would have eventually found its way to the discount rack and then the dumpster. At least this way we're getting some money in that we can use to buy actual cool books/collections, pay salaries and bills, etc.
So... something like Neutron Radiography: Proceedings of the First World Conference San Diego, California, U.S.A. December 7–10, 1981 - https://www.amazon.com/Neutron-Radiography-Proceedings-Confe...
I make no claims that that's a such a book that's been ordered, but that's the type of book that the reddit post references.
You seem perfectly fine living in a world where your flea market is devoid of books.
Also, Anna's text implies that the books are scanned and then mischievously destroyed so nobody has access again to the content. That's not the case: the books are "destroyed" before scanning, by dissasembling them in pages so they can be feed to the scanner. Scanning while keeping the book intact is difficult, as you need to software-unwarp the page before OCR'ing it, and expensive as you either need specialized scanners or humans doing it.
Yes, there are plenty of books, many were printed, many have lasted a very long time (plenty over 100 years!).
That says more about the success and utility of the technology than it does about whether individual books should be shredded.
These comparisons are starting to get ridiculous. Why are so many people assuming there is exactly one copy of all of these important books available, that it’s sitting in the warehouse of a bulk book reseller, and that Anthropic is destroying the lone copy?
Your local library throws out books every year and nobody thought twice about it.
After being clickbaited and ragebaited, media consumers feel deeply anxious and angry. But, explaining that they are angry over a boring situation feels silly, not righteous. So, they give summaries, impressions, sometimes extrapolation of the bait they have been consuming. That feels righteous.
This observation applies to a wide variety of topics trending in the various media every day. Distinguishing injustice from ragebait unfortunately requires non-trivial effort from the reader.
Why are you assuming that each book gets scanned exactly one time and then never again? And why are you assuming that out-of-print books remain easily accessible so long as not every copy has been destroyed?
>Your local library throws out books every year and nobody thought twice about it.
When they're damaged beyond hope of repair from decades of wear. As for books that haven't fallen apart from being used as intended and merely are no longer desired, my local library generally sells them off at bargain prices. Of course I can't speak to your local library.
I even read on a web site of a librarian that their library had stopped accepting donations because "patrons should know how to throw away their own trash".
What happens when the books don't sell?
Many books aren't lent and not bought and most libraries have limited space to store such books, thus they go where old paper goes.
Of course some rarely lent books are important and for the one person asking for it in ten years really valuable, but many still have to go.
I know it comes at a shock but truly most books are absolutely worthless.
Not all rare books are valuable. Someone’s self-published junk sitting in the garage is NOT analogous to the library of Alexandria.
Many, most, maybe all of these “rare” books are being scanned instead of just being recycled.
Not a big Reddit fan but there was a great post there from someone in the book industry talking about how non-industry people often give this great moral weight to ever book in a way that is totally disconnected from reality.
All I’ve read, as far as sources go, is a number of rare book sellers saying they’ve had a big uptick in huge orders with no price haggling. Apparently that’s peculiar. And some of them seemed a little concerned.
Now I’m certain they’re not chopping up Davincis notebooks, but I’m not certain there aren’t some that would make people wince.
And I don’t have any reason to think some reddit librarian knows what’s going on, if anything, either way.
Then they should list the names of these books otherwise I say they're alarmist.
It’s very possible I’ve just missed deeper reporting, obviously.
But otherwise I agree. I’m neither losing sleep over it or just trusting that these (historically kinda scummy) businesses are actually behaving.
If there’s a serious problem I’d like to see something more concrete. Same for hand-waving the question.
How can anyone say this with a straight face. The knowledge is not destroyed, it is transformed. You can make use of it today in the form of LLMs and the scans still exist. Nothing was lost. It's literally no different from them buying books and stocking them in a private library not open to the public. It's not called the Scanning of Alexandria because if it was, it wouldn't have made a blip in the history, Alexandria's libraries were burned, those books, that knowledge was destroyed. Then only thing being destroyed here is physical copy (again for the people in the back: a copy).
> Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available.
Those same copyright restrictions are exactly what would prevent them from sharing the archives. Your beef is with copyright, not the AI companies who are (in this one, rare, instance) following copyright laws/rules.
Hyperbole much?
Does the fact that they're being converted to an immutable digital permanent record for all time mean anything to you? Because as far as I know, the works lost to the Library of Alexandria were wiped out of existence, not simply transformed into a more durable form!
If a company is scanning material protected by copyright, it has to send a digital copy of the scanned material to Library of Congress within 5 working days.
Copyright should really be amended so that once out of print and a grace period it’s free use. I am probably more of an anarchist in this regard. Similar to my belief that anyone should be able to ingest any data you put online, once a book is no longer being print it should be able to be used for commercial or personal use for free. Similar to a generic drugs.
There is far too much garbage that gets published, let the collective hive mind figure out what is valuable.
That's what governments are for.
Sure some governments and opinions would say so but you’re making a statement of zero impact. Fix the underlying copyright laws don’t create more rules.
How does this help anything, except create more work to throw in the trash?
It’s called “mandatory deposit”
Looked into this a decade ago for publishing eBooks via my personal corp when eReaders and ePub were starting to hit big in the mainstream.
There was a time of course when you could pull my books down from archive.org as PDFs. Perhaps that time will come again.
I'll look into libgen.
> The print original was destroyed. One replaced the other.
So that there was still only one “copy” of the book. This is in compliance with the DMCA. You can make a personal digital copy of a work but then you cannot resell the hard copy and keep the digital one. Same principle applies here.
This isn't theoretical, AI companies have finished lawsuits about this and this was the ruling.
A person can digitize their own books without destroying the original. So can Anthropic. They are choosing to destroy the books for easier scanning and trying to palm off the blame for it.
I'm old enough to have been around when DCMA legislation was under discussion. Many people were dead-set against it and raised concerns over matters exactly like this. In Rainbows End (2006), Vernor Vinge wrote about a similar scenario where a robot went through the university library shredding books, and scanned the shredded pieces to recombined them into a digital archive.
Anthropic may be doing shady things and may have even done this on their own recognizance, we just don't know. As it stand, this is 100% a consequence of US copyright law, much of which was written by large corporations to protect their own assets.
You’re missing the point. It doesn’t require that they destroy the book, but it precludes them from giving it away. It’s their property, so they can choose to store it, but that has real, ongoing cost and may eventually leave unusable books anyway due to fire, pests, water damage, etc. if they’re not maintained properly.
To legally retain these scans, you must own the original book. You can sell or give away the original book (sans binding) but it's legally dubious as to whether the scan can be transferred along with it. So if Anthropic has no interest in storing thousands of loose leaf books, they are likely destroying both the original and scan as soon as possible.
At the end of the day, the only thing of value Anthropic has is the trained model which is definitely transferable.
That's not their goal or else they wouldn't be burning books. Their goal is making money no matter the cost to the society.
Look, others talked about how this fetishising of paper books is quite silly (though I don't like it when the destructive scanning is just so one could feed it into a chatbot) but I have to say, all I can do after reading the above sentence is laughing bitterly. Anthropic is an American for-profit company, any talk about "working towards the benefit of humanity" is just marketing lies, and it's always incredible to see people treat those seriously.
_Of course_ Anthropic does that.
Aren't the copyright laws forcing them to do this the very ones that would make such archives illegal? The books that could be in such an archive are the books that don't need to be destroyed.
They literally lie, cheat and steal at any opportunity they have and in any way that they think of. Do not trust a single thing that they say; it is a fool's folly to do so.
A lot of this can already be said about a lot of companies, especially almost any large company, but it goes doubly if not triply so for this new breed of company now.
Are you referring to the burning of the Serapeum in AD 391 or the warehouse fires in 48 BC?
Also, helps to clarify what exactly the commenter was referring to and possibly help distinguish the centuries-spanning decline of the Library of Alexandria from the violent fate of the Serapeum.
(But my knowledge of Alexandria extends only to episodes of "COSMOS" and "Connections").
Come on, we did without AI for all our history, we could live without it with no issue (as to me we could live without smartphones, internet, etc if we want).
I’m reminded of the screeds about the dangers of novels.
They're destroying one (1) copy of a mass-produced item for each AI company.
Public libraries destroy millions more yearly as a matter of routine.
This is just part of a CCP-aligned moral panic, along with the water use nonsense, and similar with the soviet-aligned moral panic that destroyed the civil nuclear industry 40 years ago.
Notably since all the controversy many data centers are very loud about being closed loop and with significant consideration given to other local impacts as well.
They don't have to, and the vast majority don't.
The lie and moral panic is that all datacenters necessarily waste precious drinking water; it's patently false and used by agitators to push, unwittingly or not, a Chinese Communist Party agenda.
But I understand blaming your energy waifu for its own failure is unacceptable for nuclear bros.
What was nonsense about water use?
The claim never made sense to me either, I can only assume those that regurgitated such claims never worked with HPC or even general datacenters before.
Was recently talking to a (non-technical) friend about this, she was surprised after talking about the "insane water use for AI datacenters" when I responded that open-loop cooling is pretty rare for a datacenter and I've never actually seen it used before, versus closed-loop (or just regular air-based cooling) which has no real noticable water consumption.
Order of magnitude more water is lost from wasted irrigation (e.g. during rain, of fallow fields, sprayed into windy air, etc) than data centers.
2. Datacenters have been shown to reduce utility prices. They provide suppliers with previsible long term demand which allows for cost-effective network and production planning.
There are none.
Let's view it realistically here: AI companies are parasites. Them destroying books to dumb down mankind, absolutely fits into the destruction of the library of Alexandria.
Having said that, I think the day of physical hardcopy of books, is not necessarily over, but will be heavily complemented via digital storage. For instance I only keep books that I may re-read later or read many more times, e. g. thick science books. Many other books I can keep as .pdf file without a problem.
But the problem is real. Even if these books aren't needed by anyone right now, them digitization in a single copy that end up behind seven locks at a corporation is not great, because AI doesn't replace the original. You can't to ask a neural net to give you a exact copy of a page from that book. So yeah, the post dramatizes a bit, but the point are valid. We need open digital archives.
That doesn't mean I'm against that narrative; I'm a big supporter of shadow libraries and scanning every unscanned book. But connecting to the LLM narrative here seems opportunistic and populistic.
The previously struggling second hand bookseller in my town has upgraded their car from a 15 year old hatchback Renault to a brand new Range Rover. Some Canadian company has been buying any book he can provide them for the last year. Their quotes aren't by number of books, or even weight, but by volume. As in, they pay him by the shipping container, and he sends several of those a month.
I think it's reasonable to assume the books in this supply chain which aren't destroyed in digitisation are just pulped and sold to Procter & Gamble for toilet paper manufacturing. I can't see any other fate for Anthropic's second and third copies of The twelfth edition of Vera Lynn's 1980's memoir "We'll Meet Again".
I doubt that corporations of this scale do that extra kind of… book-keeping. They more than likely buy books by the pound.
At some point they’ll hit on data center water consumption as yet another reason to support the cause.
> Anthropic spent many millions of dollars to purchase millions of print books, often in used condition. Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals. Each print book resulted in a PDF copy containing images of the scanned pages with machine-readable text (including front and back cover scans for softcover books). Anthropic created its own catalog of bibliographic metadata for the books it was acquiring. It acquired copies of millions of books, including of all works at issue for all Authors. Anthropic may have copied portions of Authors’ books on other occasions, too — such as while copying book reviews, academic papers, internet blogposts, or the like for its central library. And, Anthropic’s scanning service providers may have copied Authors’ print books along the way to delivering the final digital copies to Anthropic. But neither side here specifically raises legal issues implicated by any such copies. Nor will this order
Also the summary:
> To summarize the analysis that now follows, the use of the books at issue to train Claude and its precursors was exceedingly transformative and was a fair use under Section 107 of the Copyright Act. And, the digitization of the books purchased in print form by Anthropic was also a fair use but not for the same reason as applies to the training copies. Instead, it was a fair use because all Anthropic did was replace the print copies it had purchased for its central library with more convenient space-saving and searchable digital copies for its central library — without adding new copies, creating new works, or redistributing existing copies.However, Anthropic had no entitlement to use pirated copies for its central library. Creating a permanent, general-purpose library was not itself a fair use excusing Anthropic’s piracy.
https://copyrightalliance.org/wp-content/uploads/2025/06/Bar...
Just publicize a rare books vault where you put the older editions that aren’t in a lot of library catalogs. Use non-destructive scanning for those.
Align yourself with the image of safeguarding something. It seems like a no-brainer given various themes I’ve been hearing in criticisms of these companies.
Maybe the hope was to just bury the book destruction under the rug, but the cat is out of the bag. Publicizing a state-of-the-art rare books preservation archive is now a good move.
Tech tends to love associating itself with a classical tradition or something. Name it after the library of Alexandria. It would be a huge cultural loss if that were to burn down again. Thank God for our big AI companies that keep the archive intact.
Actually, I assume it would be separate archives, since I assume there’s a something of an arms race in getting training data that competitors don’t have, but really, who would complain that there are multiple archives? That sounds like a good thing. And what big AI company would want to be the odd one out for not running an archive?
A book vault would still be useful for out-of-copyright works, but this would only cover a (probably relatively small) portion. Also, I'm not sure how easy it is to reliably determine copyright at scale, so they might just decide that it's not worth it.
At this point my only hope is that in the long run these scans make it to the public somehow (leaks, copyright changes/expiration, whatever), where they can then be accessed and preserved by everybody. Then we could have our true digital library of Alexandria.
first they came for cookbooks, but i was no chef so i said nothing. second they came for handicraft, but i do not toil with fabrics or glue. next they came for homesteading, but i loathe the outdoors life. after, they came for biography, memoirs, and letters, but i am bored by the dead. finally, they came for my own little little interest, but nobody was left who appreciated books, so they too were ripped to shreds.
I ask "Why save physical books?"
If they are truly rare, then they are likely not valuable, otherwise there would be more copies or their contents could be found elsewhere.
Owning and storing physical books is not free, there is a real cost. Even owning and storing the scans is not free, especially when IP rights and challenges get involved, since a scanned book with no distribution is worthless.
I disagree with putting books on a pedestal and saying they must be protected. If they were any good, they would stand on their own, but if we have to run a moral crusade to save them, perhaps we are all better off if they're destroyed.
Sometimes I don't even know how to respond to comments here. I don't want to be rude, but you just have to give this a moment of thought. Is all the media that you find valuable common? I know that's not the case for me based on my own experience.
Doesn't that make them even worse?
"likely" being the keyword here, what about heavily censored books?
The problem isn't with morals or copyright. What we're up in arms about is case law. Past rulings have implied that destruction of books significantly contributes to the process being "transformative". This encourages companies to destroy the books. I think this is really dumb.
Why save physical books? It's because the reason to destroy them isn't good. If you think there's too many bad books out there, that's a different argument. Maybe your fight is against consumerism, I don't know.
That's what I keep saying about the van Goghs I burn to heat my home but everyone is still mad at me!
The Gutenberg bible is rare, its content is not rare, and some would argue that the content itself is not valuable, and yet the Gutenberg bible is valuable.
Same can be said about many books.
- Reader: narrative
- Collector: scarcity of the physical artifact
- AI Company: language samples (quantity, variety), facts
And about what topic have I heard more complaints in the last 15 years (although the latter topic is just a few months old)?
Why is that?
If you say that I'm indeed wrong, and the latter one IS indeed much more important, then please tell me why? What is wrong with me then?
After around 2000 books started coming out in digital format, so at least there are digital copies of many of those, even if they are still under copyright.
I wish people cared 25 years ago. Unwanted books in boxes are everywhere. Its a false hysteria. You can still get any book you want, digitizing is the best bet for more readership.
- Reinforcing outdated, disproved or otherwise incorrect information.
- Reinforcing outdated forms of communication e.g purple prose.
You could counter both by giving more weight to recent text and I suppose the extra data may help for tracing references and the evolution of ideas through history. If this is what they are resorting to it does feel more like "marginal gains" territory rather than ASI imminent territory
Much like the pushback we are seeing from the citizenry against things like Flock and AI datacenters, we can push _forward_ too by ensuring important institutions (legal or otherwise) remain funded
https://old.reddit.com/r/OutOfTheLoop/comments/1vszifd/what_...
I do wonder who is running PR at the major AI companies, as they seem rather insensate to how they are perceived...
It would likewise earn goodwill from a lot of people.
Berne Convention https://en.wikipedia.org/wiki/Berne_Convention (182 parties)
TRIPS Agreement https://en.wikipedia.org/wiki/TRIPS_Agreement (164 parties - part of WTO)
The United States can't make copyright weaker than what those agreements require without pulling out of the WTO.
The core of copyright law is about who has the right to redistribute a work. If I buy a print of a photograph, scan it and use that as my desktop image... I can do that. I cannot redistribute the scanned image, and if I was to sell the print later I should delete the scanned image.
Note that format shifting is covered under fair use... which is what training is taking place under. However, that doesn't mean that they can release that format shifted content... nor can then re-release the original work if they are retaining the format shifted content.
https://library.georgetown.edu/copyright/fair-use-reformatti...
Under § 106 of the Copyright Law of the United States, the owner of the copyright in a work has the exclusive right to make copies of that work, unless an exception applies. When considering reformatting media, please note that individuals do not have an automatic right to reformat a work from one format to another. In order to legally convert media, your use must fall into one of the following categories:
you own the copyright in the work,
you have permission from the owner of the copyright, or
you have done a fair use analysis and have determined that fair use appliesIf their PR teams did have a response to these actions, it would be something to the effect of:
"would you rather china destroy all the books and gatekeep the knowledge?"
* download and publish a book as an individual -> 100% lifetime jail + 10x your whole lifetime earnings/revenue - Aaron Swartz
* download and publish a book as a company -> fine 1% of revenue
* scan and publish a book as an individual -> legal issues, 100x fines of your yearly 50k donations
* scan and "publish/train" a book as a company -> okay, lets ban chinese models, they are distilling your model
Which is really why the outrage cycle over Anthropic's actions is largely misplaced.
How is this a question at all? They’re scanning books because the courts determined that it’s the only way to use that data. They are forbidden from using digital copies found on places like Anna’s Archive. They must acquire and scan the book.
They cannot redistribute the book. The Internet Archive tried that and the courts shut it down. You cannot scan a book and share it without violating copyright law.
Hence the “leaking” part.
Why on earth would they? Ultimate point for these companies is to make a ton of money, obviously they won't shoot themselves in the foot and give away whatever advantage they have, especially not to a free archive which is about doing good in the world, which probably isn't profitable enough for a company to care about.
I think books will exist but they will be written with the help of AI.
The context will still be a human mind behind the words in the book.
Clearly that makes assumptions about the function of government and the complacency of copyright holders.
some form of mandatory deposit law has existed since 1790 in the US. other countries have similar laws. there is a very good chance the government already has a copy of anything that has been copyrighted
I do not know much about Anna’s. Is it Chinese or is this just one of many worldwide helpers?
It's common practice to cut the spine & binding off a book, scan the loose pages, & drop the rubber-banded loose pages in a box somewhere. The book still exists, just without its binding.
It's not bitcoin.
Old books shouldn't simply vanish. That is history, art, authorial creation. Once the last copy gets crisped, it is lost.
Really can't wait for this outrage cycle to finally die. There's a lot AI companies are doing wrong. The wording of Project Panama is really some villain-type shit. But if you stop and think about it, this is really not concerning. Like, at all.
0. In the Anglosphere, the British Library (UK and Ireland) and the US Library of Congress (USA) preserves all books published in their territories. So if Anthropic is pulling a 1984/Fahrenheit 451, "gating access to information", they are really doing a ridiculously bad job at it. I'm pretty certain similar depositories exist for other jurisdictions.
1. Rare does not automatically mean it has some obscure knowledge nor does it mean culturally/historically significant. The oldest book the news outlets could mention is a 1970s manual on soil mechanics (Guardian link in your sources). How "obscure" is the knowledge in there, you reckon? How culturally significant is that book? Odds are, a good chunk of that book is outdated knowledge at this point, useless for anything practical other than knowing what people in the 70s thought about soil mechanics. And for the latter, there are a bunch of other soil mechanics manuals from the 1970s lying around.
2. No, despite the admittedly cartoonishly villainous description of Project Panama, Anthropic is not out to get every single manual on soil mechanics from the 1970s out there. They just need to cover enough subject breadth in their training data set. Getting redundant copies is a waste of money. See also, point 0.
3. You argue a lot for the nostalgia and sentimentality of old books, which, okay, but that is hardly beside the point. Libraries and publishers destroy books at a greater scale than Project Panama when there's not enough demand for them. That's not an affront to history or culture just the consequence of economics (see also, point 0). Similarly, y'all outraged about old and second hand books being destroyed but odds are, they would've been trashed anyway even if Anthropic didn't get their hands on them. We've had all the time in the world for someone else to buy them to keep them from big bad Anthropic but no one did. I'm just happy for the booksellers who made some money out of this AI-infected timeline.
4. It's not like Anthropic is buying books from Dr. James Justin Sledge---books not covered by point 0 in other words---and chopping that up. The fact that one of their reported suppliers is "ISBNdb" should clue you in that they are interested in books from the ISBN era (and hence covered by point 0). Fact is, it's a terrible investment to "gate knowledge" from the really rare and old and culturally significant books. They are über expensive and for what? Even more useless is the AI trained on those books! Imagine asking ChatGPT how to cook a large squid and it answers in the style of Thomas Hobbes.
I do think book copyright law needs a ton of work but I think most folks are simply taking their bias against AI and creating hyperbolic scenarios. I am sure there are some gems in the lot and I am equally certain they may be scanning dupes of the same material but even at scale I have a hard time seeing the significance. Most of these published work in the last 60 years is absolutely junk garbage. The good stuff usually has a lot longer run so more volume in circulation. You can go pick up lots of books that are 100+ years old for a couple bucks or cheaper because this stuff has no value.
LLMs have topped out in terms of language fluency. You’re not going to get a smarter model with 250 trillion tokens than with 25 trillion tokens. There are still other gains to be made in the LLM/LRM space, but they don’t require ripping up rare books.
And they’re doing it destructively because it’s cheaper. That’s it. They absolutely could scan nondestructively. They’re trillion-dollar companies, and they do this in a shitty way to save pennies.
The insanity of attempting to prevent AI learning (which is a direct consequence of the nature of observable information) because of the myth of intellectual property is the main driver of this type of behavior.
Nevermind that they do that because those books aren't being borrowed and no one wants to buy them. Destroying books to destroy knowledge is bad, destroying books because no one wants them or there are ample copies, _especially_ when those books go on to live forever as a digital copy, is a completely different story.
It boggles my mind that of all places, here on HN, people can't understand that. The knowledge is not destroyed, only the physical vessel.
I wasn't even making an unusual or particularly strong statement
Im going to try and be patient and explain my position. I just feel that actions that are destructive should always be seen through a lens of suspicion.
IE, lets destroy this wood to reduce the chance of wildfires. Ok, well why that wood? we sure there arent other interests at play? Ie a carpark wanting to be built by a local authority?
As you become older and witness large acts occur due to unclear scheming, you start to become sensitive towards destructive acts, feeling there should be a lens cast over it.
Maybe Im just paranoid.
Sure I can understand there must be tons of books that are not needed, but books, are object that include native function that IN the right context becomes priceless and hence should be treated with a degree of concern.
Yes destroying a book is in my opinion identical to a book burning.
One day there will be no hard copies, the digital books are hence vulnerable to patches, whether its due to a new political movement or a sudo abled hamster running on a keyboard... and if that occurs information will be permanently lost.
I would have thought a HN user understand well the importance of backups.
I dont want to sound mean, perhaps you also support keeping some hard copies as backups, but Im just explaining my concern.
To be honest I think everything I am saying is just common sense, I was just making a joke earlier about the architects deciding to do this because they had a bad tuesday.
I'll admit I have less concern with (every) digital copy being altered and if we assume a future where all computing is completely locked down and governments have the ability to reach in a tweak anything then yes, we are screwed. But we are screwed if we allow that to happen even if some of us have squirreled away physical books. In that techno-hell (of completely locked down computing with no open options) then some of use will have "illegal" computers that still do what we want and I see that as no different than keeping physical copies.
Put simply: Preserving the original does not require a physical substrate (or at least one made of paper, obviously computers run on physical hardware).
They have done zero to destroy the durability. In fact, it’s probably more durable.
If the physical copies are scarce, that is due to the publisher and copyright laws and not them buying and converting a single copy.
Agree. While it certainly isn't the most environmentally friendly to render huge stacks of paper into waste, the real issue is copyright creating scarcity (inability to copy the thing).
Are they? Or are they just using it for training and not keeping the digital copy afterwards? And even if they are keeping a digital copy, does that actually matter if they never release it?
You’re making a distinction here, but training is something you repeat for every new point release, so you need to keep the data if you want to use it for training.
As for long term benefits, it could one day be sold as a service, once copyrights have expired on the works. We can't see it today, but that is purely the result of the law and what the law intended to do from the start, you don't see a copy unless you pay for your own.
I’ll leave
Heck even YouTube should be legally obligated to preserve videos, given how they now hold the largest visual documentation of human history.
Companies are paying for the books now, great! And destroying a book is really the best way to scan it. I could destroy the books I buy myself if I wanted, they're my books and I can do whatever I want with them. A book is really just a stack of paper, I don't get the sacred feeling attached to it. Especially since books are printed in thousands to millions of identical copies nowadays.
Now what books exist in a single copy that destroying it would amount to losing knowledge? In any case, I would also expect such books to be very old, have very little useful content for training to start with, and be expensive enough to buy to make training on them unprofitable.
So what's the problem here exactly?
Also from the article:
> It’s outrageous is that it’s legally permissible, but ethically, it’s an extremely serious crime against humanity.
I'm all love for Anna's archive, but still, I find it rich that they now find themselves in position to make strong ethical claims.
If you've read the articles covering this issue, you'll be aware the concern is over the fate of rare and out of print books, rather than your straw man (ie those available in 'thousands to millions' of copies).
https://www.bbc.com/news/articles/cp3rprx2wl4o
> "A recent academic text published in only 100 copies, 75 of which are already in libraries, may be very rare on the market - but it is perhaps not such a great loss if one copy is destroyed," says Derek Walker, owner of Edinburgh bookshop McNaughtan's.
> "But we have, and have sold, books which are for example the only known surviving example of an edition from the 18th century.
> "It would be a much more significant problem if one like that were to be bought for destruction, having survived this long."
> "It would be a much more significant problem IF one like that were to be bought for destruction, having survived this long."
I agree with the sentiment of course but it is really a huge IF they are doing that.
IF they wanted to train on, say, Leviathan by Thomas Hobbes, why buy an expensive edition from the 1600s when they would get the same text from a Penguin edition for a fraction of the price? It gets much cheaper secondhand too of course.
I'll go further, why would they want to train on expensive rare and out of print books? Are they, perhaps, competing on an AI benchmark based on extensive medieval knowledge of the cosmos? There's been a lot of pearl-clutching about lost obscure knowledge but y'all really reckon that kind of knowledge is valuable to LLMs?
This idea that there can only be merit in a work if it's commercially viable is incredibly ignorant, and if that becomes the standard for whether a work is preserved or not, we stand to lose a great deal of our cultural heritage.
Detestable if they're doing it anyway to prevent competitors getting hold of it.
There are multiple instances of a book being referred within another book, while at the time the author had access to it, we might not have it today. Taking a rare book and destroying absolutely erases that link we have with the past.
You not seeing a problem with this is the core issue, it's probably why the people doing it (it's people destroying these books not aliens) just shrug and don't feel too bad doing that.
Old books are even more crucial than today's books due to how uncommon it was to have something written/printed and bound. Many unique and single copy books explain to us a ton of things about the past, sometimes for funsies and sometimes for useful findings. Destroying old books is akin to destroying the closest we got to time machines.
OR in-print modern books that they can get for cheaper by buying used. The whole thing is a manufactured outrage over something that doesn't matter. They aren't breaking into museums to steal their only copy of a book and burn it. They are digitizing books. If anything I commend them for what they are doing. If the alternatives were that book rotting on a shelf or being thrown away they doing a great service preserving it, even if they don't make it available publically (which they can't for copyright reasons). It's literally no different from them stocking a private library with these books, except it's better because digital copies are much easier to preserve.
It’s powerful symbolism. It reminds me of that tone-deaf iPad ad that sparked outrage in 2024. The one where all the cultural artifacts were crushed in an industrial press to make a soulless slab of glass. And that was Apple, who is generally well-liked by the public.
Of course you can. Nobody is saying these companies aren't allowed to do what they're doing.
But what they're doing is disgusting and something I can't forgive. It's an escalation of the attacks against society that these companies have been engaging in from the beginning. This isn't about legality, this is about what's right.
Isn't this just doing the work for the AI companies??? Then the AI companies can simply download a copy of Anna's Archive.
Google wanted to share the whole of Google Books 15 years ago, too, but they were sued to hell, so now you get a watered down search functionality.
Baloney. They aren't forced to destroy books. They could leave well enough alone and not scan them to begin with. They're voluntarily choosing to do this.
That's rather beside the point. I don't think anyone is arguing that what they're doing is in some way illegal.
> no modern book
I'm less concerned about modern books.
I'm not convinced they destroy the books to obey the law, lol.
This is such dangerous train of thought, to give them the benefit of being forced to destroy books. Why is that exactly, and who is forcing them? You can also, you know, find another way?
Like the data centers who currently use very dirty energy acquisition methods (not all of them), are they also "forced" to do this, because they too need to make as much money as the other ones? How long would you continue this idea of others "forcing" for-profit companies to try to make more money, regardless of consequences?
Destroying books used to be an obvious dumb, stupid and shit idea, not sure how somehow a for-profit company making of a digital copy for themselves of the book before destroying it, suddenly makes it not a shit idea for the rest of humanity.
I don't blame the companies for either wanting digitised works, nor following the law. I blame the absurd court ruling, and I blame the publishers for not having digitised the old works themselves, in which case they could just sell e-books to the labs... They are after all the only ones who can legally do it non-destructively.
A. Company destroys a book forever for training. Its scan is locked away forever in company records. In this case:
- This information is locked away in the improvisations of an LLM. It is no longer possible to directly access the information as written by the human being that authored it. This constitutes the loss of literary history, or at least loss of access to that history to the general public.
- The price of the book is no longer distinct from the general price of "inference". It becomes increasingly impossible to pay for specific information, instead you are charged by the meter for general machine inference, which doesn't even give you access to a specific text.
- The provenance of information is totally destroyed. This causes potentially unresolvable problems of authority and citation. If the original source is lost, how are we to know if a random LLM claim about an obscure topic or specific niche text is even true or just hallucinated?
B: Company uses freely available scanned copy of the text:
None of the issues above obtain, since anyone can still access the actual book. Most importantly, this reduces the power companies have to force everyone to continually pay for a derivative form of the book's information in perpetuity in the form of token costs.
I much prefer B.
Because these companies have no real strategy beyond trying to capture any and all information they possibly can to try and lock it away and charge the public for it in perpetuity.
The rhetoric on this topic is reminiscent of the rhetoric regarding data centers: some noxious combination of misinformation, misunderstanding, and sensationalism, wielded against technological progress.
Using a few gas generators for a transitional period isn't great either, but effect of those is very temporary. I fear permanence in this individual deal with the devil made for speed and cost.
Here's another question: The exact reasoning for copyright is made clear in the constitution: "[the United States Congress shall have power] To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries."
What courts changed is failing to promote the progress of science and useful arts by destroying knowledge/art/books and access to those books, supposedly with the goal of maintaining market demand for those books. I mean this reasoning is so bad, so warped it could be used as James Bond villain humor.
Courts have failed to provide authors and inventors with exclusive rights, in fact they have destroyed a right authors effectively had. This decision goes against both the spirit and letter of the law, all because billionaires don't want to respect copyright anymore.
If you're going to do this, why have copyright at all? Can someone explain to me HOW you can explain that the constitution still supports this system?
But, of course, it gets worse. EU courts have decided the opposite, namely that without the author's permission you are not allowed to train an ML model on their texts. Ie. this becomes an additional right authors have, an additional thing authors can license (or not). Now that SOUNDS good, and you can bet the EU commission will be publishing about that. But it isn't good.
Of course the EU has demonstrated their usual do-nothing attitude. They have stated they are not going to do anything about people outside of the EU blatantly violating EU law, and let them profit inside the EU of violations of EU law, thereby destroying authors' income. As to the question why there is any need for EU law if you're not going to act on law violations ... no answer on that front.
To put it differently: why is chatgpt.com, claude.ai, gemini.google.com, ... not banned across the EU? Why are payments involving violating models allowed to go through, given that EU courts have sided with authors? What is the point of having EU laws at all?
But it's far worse: the EU commission is attempting to make their employees use a US model (chatGPT [2]), in violation of EU law, internally in their own organizations. Now from what I hear, they're failing at making people use it, WHILE paying US companies for illegal models.
I mean I hate what the US government has done, but the EU is far worse. They are officials, they are the institutions ... and their public claim to defend authors, their court judgements, their public stance ... is just an outright lie. I mean how else can you call this? 90% of the people involved here are lawyers, from the very top to the bottom rungs, all overwhelmingly lawyers. They know they are going against their own law, and doing it anyway. That is, at best, lying. The EU commission, even the courts and the EU's own bureaucracy will not follow the law internally, NOR are they making anyone else follow EU law!
The real effect of EU decisions: only Mistral, and other EU model providers are forbidden from, and punished for training on copyrighted data without permission. Everyone outside of the EU can just do it without permission and will not face any kind of consequences for that in the EU. EU companies (ie. hugging face) are hosting, for free, models in blatant violation of EU law.
Which is even worse than the US government stance in my opinion. EU has directly chosen for the worst possible of all combinations:
a) EU companies making ML models have to self-sabotage against their competition.
b) EU authors receive ZERO protection from the law. Not because the law doesn't support their case, but because the institutions whose only reason for existence is to enforce the law won't do their job. In fact THEY THEMSELVES violate EU authors rights.
Obviously, under these circumstances, AI model companies are going to outcompete musicians, authors, even movie studios. It's only a matter of time. And that is not a given, it is an explicit choice both the US and EU governments are making.
[1] https://commission.europa.eu/document/download/f0b8d4c3-51aa...
[2] in their source you can see what models they were likely using internally 2 years ago: https://github.com/openeuropa/gpt-at-ec-php-client
Overall , smells like propaganda. How many books on earth you can openly buy that has practical “knowledge” value , not just historical one , and exist only in one physical implementation ( book ) . There were numbers of initiatives for a decade to digitalize all valuable knowledge.
It seems the initiative was secretive for a different reason . Buying a single book not allows you to distribute the content of the book , hence 1.5 billions fines.
Usually I personally not on copyright people side , but in this case you clearly see the system functioning. Know knows how the book business will look like in 10 years , but now book publishers doing their job by brining ai companies to court