ed: to clarify, I totally agree that a huge chunk of the value in LLMs is coming from the source material. My point was just that training an LLM takes more resources and expertise than distilling from an existing LLM so I don't think the equivalence between training and distilling is entirely justified.
It's the most CS-major take ever!
The LLM output, is not the same as the input - there is value add.
Of course works used as raw inputs to LLMs required work and are reasonably subject to IP concerns - but they are different.
It's possible that the LLM makers 'owe' the content creators that created the content they used to make their products - it's an interesting but separate question.
We could very well end up where content IP is protected, LLM output is not and visa versa with reasonable legal founding, doubtful but plausible.
But it's debatable if that's the case.
Google stores copyrighted content and produces in in their product.
Also - it's fair game to use snippets of things here and there, if the derived work is novel, which I think it is for LLMs, mostly.
I do agree though, that we ought to draw the line somehow.
How, and why?
> We could very well end up where content IP is protected, LLM output is not and visa versa with reasonable legal founding, doubtful but plausible.
That is the current state of legal rulings - LLM output is public domain, not copyrightable.
Our current laws simply weren’t built for this and I expect the legal status of LLM output is not going to be resolved until Congress actually legislates on this topic.
How, and why?"
How are they even remotely the same?
They're not even used the same way.
One is raw data input, the other is training content - designed to train LLMs.
One is a set of IP derived for other purposes entirely, and has esablished IP law - how you can use someone else's creative work or not ... for LLM outputs, less clear.
Or maybe they're going through an intermediary "transfer station" that's breaking terms of service:
https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens...
Anthropic's own copyright infringement could apparently be forgiven for 1.5B USD after all, so maybe there's a price that breaking the distillation clause for is acceptable too. Or some other arrangement.
Like look, I'm not a native speaker, sure. But I think when someone says "value add", that means there was value there (which you claim they're rhetorically erasing), and then that was added to. Under no interpretation of this phrase do I get an erasure of prior value.
So certainly, as long as words mean anything, no, they absolutely did not say or suggest what you claim they did, and what you extract a thus unreasonable amount of obnoxious schadenfreude from, while throwing in an insult for funsies at the end.
It's the second time I feel compelled to reach for this just today: https://i.kym-cdn.com/photos/images/original/002/659/979/108...
Writing books, building Wikipedia, and answering questions on online forums takes a lot of resources and expertise that scraping didn't. So at the very least, we're already one rung down the "maybe you should've asked" ladder.
This is not automatically true. Training and distillation use the same underlying infra and method and there is no intrinsic differences in between.
It took me a year to write a book. It took OpenAI and Anthropic a fraction of a second to ingest it. Do you understand now why I give zero shits if it takes Anthropic a billion to train a model, and Moonshot 10k in API cost to distill it?
If that is the whole point you need to clarify why this is the case on an objective level.
I would say building a comparable model using any means necessary (just like what Anthropic and OAI did) at a lower cost is actually more valuable to soceity and Monshoot is arguably generating more value with less.
If the distilled model is cheaper, then it's just LLM's getting LLM'ed.
Raw materials vs. Value add.
They are different things, like ore and metal.
Distillation is a new thing we need to understand, it's probably closer to IP than not.
We're talking about things like text people wrote, not some kind of raw data floating out in the ether.
Ore has value, a different kind of value than the output of the refinery.
https://arxiv.org/abs/1503.02531
although i doubt there has been a legal case over it yet in the context of the legality of stealing shit but IANAL.
It's completey insane that we still don't know how Open Source would work, that the laws are vague and we're still technically waiting for the courts to decide on cases.
The government should a) legislate and b) create test cases and run them through the courts so that we can have clarity.
That would stun me, but it's a little hard to read.
The recent announcement that AI-assisted research produced a counterexample to the Jacobian conjecture--a long-standing open problem in algebraic geometry--shows the original value AI can create. The result was not copied from a textbook; it emerged from AI learning from existing material, much as a human does, and then applying that knowledge in a new way. If that's a violation of copyright, then a human doing the exact same thing would be a copyright violation too. But it isn't.
True, if the human's access to the book was legal
A great deal of training was on the open web, no one should complain.
But at least Meta and Anthropic were caught red handed taking copyrighted works, illegally, for training
I think international IP laws are too strick and onerous, but they were broken to train these models
It’s massive copyright infringement.
The human buys the books.
I wouldn't want to live in a world where technology or general people's wellbeing was held back by obsolete laws that ended up lingering on just to protect undeserving special people at the expense of the rest of society. Remember guilds for tradesmen? They were also a monopoly given by the government to special people. They had their purpose but nowadays we have different ways to keep tradesmen working effectively like license requirements and insurance.
Just to be clear, I think we do still need copyright, but that we might be in a transition period where it has to be redesigned to adapt to AI.