upvote
Anthropic was asked to pay book authors because it trained on pirated downloaded books.

What about books and art where the author/artist does not authorise AI to train on it? They do happily train on it, ignoring their "ToS".

This is just double standards, a slap on the wrist to not worsen the situation with authors imo.

You could also say the Chinese companies are doing the same - they _do_ pay for their Anthropic subscriptions after all.

reply
>Anthropic was asked to pay book authors because it trained on pirated downloaded books.

No, it was asked to pay book authors because it pirated copies of books and stored them on their hard drives. The ruling had nothing to do with training.

reply
Thanks - I didn't verify this personally, but still stand corrected.

If training wasn't considered outside the law, this goes on to make the point about double standards for US vs Chinese model training methods.

reply
Pirating millions of dollars worth of books would land people in jail. Government went after Aaron Scwartz for "pirating" scientific journals.
reply
Aaron Schwartz was re-publishing the journals, which is the core concern or copyright and indeed where the word "copyright" comes from. Anthropic could get in a lot more trouble if their models are caught reproducing copyrighted work from their training set wholesale.

I think the Aaron Schwartz case is incredibly vexing because he was obviously acting out of a sense of altruism without personal self-interest. I don't think he deserved the book getting thrown at him like that. But the whole copyright system, which people seem to think is simultaneously good and bad, kinda rests on not allowing those kinds of violations

reply
I've linked this several times, but LLMs are capable of reproducing entire books. Researchers were able to extract books nearly verbatim: https://arxiv.org/abs/2601.02671
reply
I didn't think about this parallel, but that's just so sad.
reply
deleted
reply
The judge explicitly ruled that training on the books was fine.

They just should have bought them, rather than pirating them.

Also LLM output is not IP (in itself) in the first place, nor would Anthropic want to claim it is and that they have rights to it - that would drive paying customers away.

The issue comes down to at most ToS violations.

reply
> They just should have bought them, rather than pirating them.

Bought, scanned and destroyed them I believe. The judge okay'd Destructive Scanning.

reply
Destruction is not necessary. Google Books is a solid precedent. You can keep the content, you just can't make significant parts publicly available
reply
> The court also held that the third factor favored fair use as to the purchased library copies converted from print to digital because the purpose of the copying was to keep the books in its library but with more favorable storage and searchability properties. This purpose required copying, there was no surplus copying and the source copy was destroyed. With respect to the pirated copies, however, the court held that because “Anthropic lacked any entitlement to hold those copies” and retained them “even after deciding it would not make further copies from them for training,” this third factor weighed against Anthropic for that particular use.

https://www.loeb.com/en/insights/publications/2025/07/bartz-...

You need to destroy the _physical copy_ that you scanned.

reply
The settlement doesn't cover the actual training on copy written material
reply
That will never happen. Copyright is not a thing in China.
reply