upvote
No, they settled that yesterday, so all is forgotten. Press releases were queued for today so just in the nick of time.
reply
While I agree on a moral level, I think there is a distinction to be made. Training a SOTA model takes a huge amount of resources and expertise so the people doing the training are adding a lot of value along the way. I think this is much less true for distillation (which is kind of the whole point).

ed: to clarify, I totally agree that a huge chunk of the value in LLMs is coming from the source material. My point was just that training an LLM takes more resources and expertise than distilling from an existing LLM so I don't think the equivalence between training and distilling is entirely justified.

reply
I like this comment because its argument only makes sense if you assume that the entire world's output of books and art did not require a huge amount of resources and expertise to make, nor did it add any value.

It's the most CS-major take ever!

reply
If turning other peoples copyrighted work into a model is transformative enough to be protected then so is distilling that model into a different, better, model.
reply
The models were built using copyrighted works, so why can't models be built using other models?
reply
This is a misrepresentation though.

The LLM output, is not the same as the input - there is value add.

Of course works used as raw inputs to LLMs required work and are reasonably subject to IP concerns - but they are different.

It's possible that the LLM makers 'owe' the content creators that created the content they used to make their products - it's an interesting but separate question.

We could very well end up where content IP is protected, LLM output is not and visa versa with reasonable legal founding, doubtful but plausible.

reply
Lossly storing IP in LLM itself, and using IP for training (so it’s lossly stored in LLM), without licensing these works or otherwise following license agreements (eg GPL) is infringement. Using then this product for commercial activity is a smoking gun.
reply
"Lossly storing IP in LLM itself, a" - that part I'm inclined to agree with.

But it's debatable if that's the case.

Google stores copyrighted content and produces in in their product.

Also - it's fair game to use snippets of things here and there, if the derived work is novel, which I think it is for LLMs, mostly.

I do agree though, that we ought to draw the line somehow.

reply
> but they are different.

How, and why?

> We could very well end up where content IP is protected, LLM output is not and visa versa with reasonable legal founding, doubtful but plausible.

That is the current state of legal rulings - LLM output is public domain, not copyrightable.

reply
This misstates the small number of legal opinions and orders on this topic, none of which form binding precedent outside the districts where the cases happened. So even if a court had found that “LLM output is public domain” (none did) that wouldn’t make it “the law” until it went up the appellate system and was upheld.

Our current laws simply weren’t built for this and I expect the legal status of LLM output is not going to be resolved until Congress actually legislates on this topic.

reply
"> but they are different.

How, and why?"

How are they even remotely the same?

They're not even used the same way.

One is raw data input, the other is training content - designed to train LLMs.

One is a set of IP derived for other purposes entirely, and has esablished IP law - how you can use someone else's creative work or not ... for LLM outputs, less clear.

reply
I don't think that's what it's saying at all. It's saying that there's a level of creativity in model creation that isn't present in distillation.
reply
Maybe, but it's not like their AI is likely to repeat it back verbatim so it's unlikely to be a copyright violation. It seems like at most, they would be breaking Anthropic's terms of service?

Or maybe they're going through an intermediary "transfer station" that's breaking terms of service:

https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens...

reply
Yes, it's just a ToS violation at present. Those are legally binding though, despite the common adage. What that really translates to here though, anyone's guess.

Anthropic's own copyright infringement could apparently be forgiven for 1.5B USD after all, so maybe there's a price that breaking the distillation clause for is acceptable too. Or some other arrangement.

reply
Yes, this is what I was getting at.
reply
No? They outright say the opposite!

Like look, I'm not a native speaker, sure. But I think when someone says "value add", that means there was value there (which you claim they're rhetorically erasing), and then that was added to. Under no interpretation of this phrase do I get an erasure of prior value.

So certainly, as long as words mean anything, no, they absolutely did not say or suggest what you claim they did, and what you extract a thus unreasonable amount of obnoxious schadenfreude from, while throwing in an insult for funsies at the end.

It's the second time I feel compelled to reach for this just today: https://i.kym-cdn.com/photos/images/original/002/659/979/108...

reply
> raining a SOTA model takes a huge amount of resources and expertise

Writing books, building Wikipedia, and answering questions on online forums takes a lot of resources and expertise that scraping didn't. So at the very least, we're already one rung down the "maybe you should've asked" ladder.

reply
I suspect that, in aggregate, all of the informational output of humanity prior to 2020 has taken more resources to produce than the last few years of LLM research.
reply
I don't know man. This reads like "yeah we stole your grain, but making bread is hard."
reply
It sure is, but it doesn't matter. Whatever position that generates more economic activity is declared legal using some nonsense retconned logic "because we said so".
reply
Probably not as much effort as writing books and creating art the models were trained on.
reply
> training an LLM takes more resources and expertise than distilling from an existing LLM

This is not automatically true. Training and distillation use the same underlying infra and method and there is no intrinsic differences in between.

reply
Yeah, there's a difference. One party spends a bunch of resources doing something illegal and extremely immoral. The other party spends little money doing something legal and morally neutral.
reply
They add value on top of other people’s work, often against licensing, and then commercialize this product, ie profiting from making a product out of other people’s IP.
reply
deleted
reply
deleted
reply
As an author, that's a genuinely disheartening thing to read.

It took me a year to write a book. It took OpenAI and Anthropic a fraction of a second to ingest it. Do you understand now why I give zero shits if it takes Anthropic a billion to train a model, and Moonshot 10k in API cost to distill it?

reply
Why is it less true for distillation? Everyone technically has access to Fable but Moonshot came up with the model. How can you objectively claim one is adding value while the other is not?

If that is the whole point you need to clarify why this is the case on an objective level.

I would say building a comparable model using any means necessary (just like what Anthropic and OAI did) at a lower cost is actually more valuable to soceity and Monshoot is arguably generating more value with less.

reply
You can argue that reverse engineering anything is as hard if not harder than engineering something. I can’t imagine distillation is any different.
reply
Distillation is objectively easier than training a model from scratch, that's why all these Chinese labs are doing it.
reply
Training a model is objectively easier than generating the sum total of human creative output prior to 2020. That's why the big labs are doing it. What's the difference here?
reply
The value of LLM's come from replacing what generated its training data.

If the distilled model is cheaper, then it's just LLM's getting LLM'ed.

reply
I'm sure it takes a lot of time and resources to plan and pull off an epic heist but it is unusual to see people like Thomas Crown being accused of creating value, as they're usually accused of committing theft.
reply
No - distillation is not data inputs.

Raw materials vs. Value add.

They are different things, like ore and metal.

Distillation is a new thing we need to understand, it's probably closer to IP than not.

reply
The "data inputs" were also, very much, somebody's "value added" IP.

We're talking about things like text people wrote, not some kind of raw data floating out in the ether.

reply
Did I say there was no value add in the inputs?

Ore has value, a different kind of value than the output of the refinery.

reply
distillation has been around for 12 years. it's not new in terms of ML techniques.

https://arxiv.org/abs/1503.02531

although i doubt there has been a legal case over it yet in the context of the legality of stealing shit but IANAL.

reply
Yes, I get that, but it's only now the issues are coming into the commons in a way that industry / society needs the regulatory clarity.

It's completey insane that we still don't know how Open Source would work, that the laws are vague and we're still technically waiting for the courts to decide on cases.

The government should a) legislate and b) create test cases and run them through the courts so that we can have clarity.

reply
Are you suggesting data input is further from IP than distillation?

That would stun me, but it's a little hard to read.

reply
Do you think writing books (and Wikipedia articles, and stack overflow articles, and github repos, and, and, and, and ...) is not a value add?? What terrible claim.
reply
This is why I don't give a shit that this is happening. It's actually kind of funny to me.
reply
Unless you’re from mainland China, you should.
reply
Why? Should we be held hostage to the whims of these companies and all the investors in the throes of AI psychosis, and let them do whatever the fuck they want, because if we don't then the economy will crash?
reply
Why so? If I'm from Europe or South America, should I hope Anthropic/Openai win?
reply
No you should not care unless you are a shareholder in one of these Ai ponzi companies. For the rest of the world Chinese companies matching and open sourcing llm's will keep 100 or so tech oligarchs taking over all the world economic output for themselves as all the idiot politician are unwilling to tax wealth.
reply
I disagree that LLM models are the product of enormous quantities of copyright infringement.

The recent announcement that AI-assisted research produced a counterexample to the Jacobian conjecture--a long-standing open problem in algebraic geometry--shows the original value AI can create. The result was not copied from a textbook; it emerged from AI learning from existing material, much as a human does, and then applying that knowledge in a new way. If that's a violation of copyright, then a human doing the exact same thing would be a copyright violation too. But it isn't.

reply
If you re-read your comment, you will find that your second paragraph is not evidence for the claim you make in your first paragraph. In fact, your first paragraph is just false.
reply
Let me explain it this way: If it is legal for a human to learn from a book, then disseminate the knowledge, then it is legal for a machine to do so. You may think this is not right because a machine does it at a much larger scale, but if so laws need to be updated. As it stands now there is no law that says if a human does X it is not a copyright violation but if a machine does the same X it is copyright violation.
reply
> If it is legal for a human to learn from a book...

True, if the human's access to the book was legal

A great deal of training was on the open web, no one should complain.

But at least Meta and Anthropic were caught red handed taking copyrighted works, illegally, for training

I think international IP laws are too strick and onerous, but they were broken to train these models

reply
Yup if a machine kills a human it's not the machine's fault; it's the human's. Humans doing the exact same thing as machines aren't 1:1.
reply
If it is legal for a human to do something then it is legal for a machine to do it too. Are there any counter examples to that?
reply
The didn’t pay for the books.

It’s massive copyright infringement.

The human buys the books.

reply
It's worth keeping in mind the purpose of copyright. It's a pragmatic tool to encourage investment in creative work for the benefit of everybody/consumers. We may be entering a time where there's less need to incentivize people to write books. At least not non-fiction books which are simply a collection of existing knowledge presented in an a way that's suitable for human readers. A lot of the value those authors provided can now be done by AI. Yes, the AI trained on their work, but now that it's here, we don't need new non-fiction authors quite as much as we used to.

I wouldn't want to live in a world where technology or general people's wellbeing was held back by obsolete laws that ended up lingering on just to protect undeserving special people at the expense of the rest of society. Remember guilds for tradesmen? They were also a monopoly given by the government to special people. They had their purpose but nowadays we have different ways to keep tradesmen working effectively like license requirements and insurance.

Just to be clear, I think we do still need copyright, but that we might be in a transition period where it has to be redesigned to adapt to AI.

reply
Did they borrow the book? If I learn from a borrowed book is that copyright infringement?
reply