upvote
$1.5 billion fine for downloading 7 million books from LibGen and other pirate torrents.

That's also the case where the judge ruled that training AI models on books could qualify as fair use, but storing millions of pirated works in a central internal library without licensing constituted copyright infringement. It will be interesting to see if courts consider training on data distilled from a model fair use. Assuming the allegation is true. Someone distilling data from a cloud-hosted model:

- Paid the model creator to use a publicly available product.

- Never copied or even had access to the model source code or weights.

- Created a derivative work based on the model's responses to their particular input.

- Trained their own model on the distilled output

That distilled output is arguably a collaborative creation because a distiller's prompts are their own unique intellectual property. So they never pirated anything. I'm struggling to see how distillation is copyright infringement. At most it seems to be a paying customer violating one of the license terms, perhaps akin to a "no commercial use of derivative works" clause. But in the case of giving away an open weight model, is it even 'commercial use'?

I guess if the distiller asserts copyright on the weights but gives them away, it's technically 'commercial' but even if they can win that argument, they're left with zero direct damages and suing for some value delta based on the alleged revenue they were deprived of. Is that delta the difference between the distilled model existing and the next best non-distilled open weight model existing? And then they have to collect damages from a portion of the revenue of third parties who commercially served that free model?

reply
Well I mean it still worked out for them because they wouldn’t have had the 1.5 billion to license before doing the training and the company exploding into a trillion dollar company?
reply