upvote
What theft? LLM output has never been copyrightable.
reply
I'll give you the benefit of the doubt. GP was referring to the use of copyrighted material to train LLMs.
reply
Which is not "theft" according to current legal precedent, and also common sense.
reply
Download a book and you are a thief

Download 1 million books and you are OpenAI

reply
What? HN has always praised Library Genesis for example.
reply
Hunting an animal would be considered ok by most, but scaling it to the point of damage is not ok according to most.
reply
It is theft under common sense.
reply
Only in the sense that you reading a book from the library also constitutes theft of knowledge. Does it?
reply
If I stole someone's private research notes and republished them loosely in my own words, they'd correctly be pissed

This was unpublished research that was stolen, and constitutes plagiarism and academic fraud by even the strictest definition

reply
>This was unpublished research that was stolen, and constitutes plagiarism and academic fraud by even the strictest definition

It was not stolen, it was willingly given.

reply
The terms of service does not dictate what constitutes plagiarism
reply
1. using copyrighted material to train LLMs is fair use, not theft

2. The topic we are dissussing concerns LLMs being trained on logs from previous LLM chats. If you're prompting a model and it spits out some unique mathematical insight, you do not have copyright on that.

reply
Nobody owns math and it is ridiculous to suggest someone should.
reply
> why enabling mass theft is so incredibly damaging to society. This is literally why we need a functional copyright system

What is interesting is that LLM's do not directly violate copyright. The settlements we have seen are for how the works were acquired (that was a copyright violation) not the use of the works.

The vectors of a book, or a paper, are not the paper. They are, for all intents, facts about the work itself, and more generally writing. You can not copyright a fact.

It also means that the weights, the things that (mostly) matter can not be copyrighted either.

reply
This is literally why we need a functional copyright system

To block progress. Got it.

reply
This is the literal opposite of progress: stealing from people genuinely creating, and stealing the money they should earn
reply
Funny, the stuff that was stolen is still there. A strange kind of theft.

Copyright maximalism is a bad look on a site called "Hacker News." Perhaps other sites beckon.

reply
> Copyright maximalism is a bad look on a site called "Hacker News." Perhaps other sites beckon.

Frankly it's more of an insult to the "hacker" name to be apologising for big companies profiting off of frontrunning existing work for PR purposes, if the claims about piggybacking on human-directed efforts/prompting are true.

Being pro-copyright in order to protect the work of an individual from being reconstituted into the corporate machine is VERY hackery. Novel use for an existing tool, to fight the dominant system.

(Of course, we're on a so-called "hacker" site hosted by a company run by squarely-establishment individuals acting in an extremely un-hackery-field (investing), so the irony here has been at least one layer deep since the start.)

reply
"Corporate machine," yadda, yadda, whatever, go sell it on Reddit. The model running on the box in my basement is almost as good as the one we're talking about here, and may in fact be just as good by this time next year... and it couldn't have existed under your proposed regime.

Yes, OpenAI is likely to be found to have acted like a slimeball in this instance, or at least the employee in question may have. But you can't fix that without making laws that will make everything else worse... and only here in the US.

reply