But it's debatable if that's the case.
Google stores copyrighted content and produces in in their product.
Also - it's fair game to use snippets of things here and there, if the derived work is novel, which I think it is for LLMs, mostly.
I do agree though, that we ought to draw the line somehow.
How, and why?
> We could very well end up where content IP is protected, LLM output is not and visa versa with reasonable legal founding, doubtful but plausible.
That is the current state of legal rulings - LLM output is public domain, not copyrightable.
Our current laws simply weren’t built for this and I expect the legal status of LLM output is not going to be resolved until Congress actually legislates on this topic.
How, and why?"
How are they even remotely the same?
They're not even used the same way.
One is raw data input, the other is training content - designed to train LLMs.
One is a set of IP derived for other purposes entirely, and has esablished IP law - how you can use someone else's creative work or not ... for LLM outputs, less clear.