Download 1 million books and you are OpenAI
This was unpublished research that was stolen, and constitutes plagiarism and academic fraud by even the strictest definition
It was not stolen, it was willingly given.
2. The topic we are dissussing concerns LLMs being trained on logs from previous LLM chats. If you're prompting a model and it spits out some unique mathematical insight, you do not have copyright on that.
What is interesting is that LLM's do not directly violate copyright. The settlements we have seen are for how the works were acquired (that was a copyright violation) not the use of the works.
The vectors of a book, or a paper, are not the paper. They are, for all intents, facts about the work itself, and more generally writing. You can not copyright a fact.
It also means that the weights, the things that (mostly) matter can not be copyrighted either.
To block progress. Got it.
Copyright maximalism is a bad look on a site called "Hacker News." Perhaps other sites beckon.
Frankly it's more of an insult to the "hacker" name to be apologising for big companies profiting off of frontrunning existing work for PR purposes, if the claims about piggybacking on human-directed efforts/prompting are true.
Being pro-copyright in order to protect the work of an individual from being reconstituted into the corporate machine is VERY hackery. Novel use for an existing tool, to fight the dominant system.
(Of course, we're on a so-called "hacker" site hosted by a company run by squarely-establishment individuals acting in an extremely un-hackery-field (investing), so the irony here has been at least one layer deep since the start.)
Yes, OpenAI is likely to be found to have acted like a slimeball in this instance, or at least the employee in question may have. But you can't fix that without making laws that will make everything else worse... and only here in the US.