upvote
> it's only now the issues are coming into the commons in a way that industry / society needs the regulatory clarity.

you mean like the regulatory clarity surrounding stealing shit to make the LLMs in the first place?

> [There is] extensive litigation on the limits of Fair Use to AI development. Currently, we only have 3 first instance decisions out of the 53 cases being tried. It will likely take a decade before we understand how Fair Use applies to any one step in AI training, let alone all.

https://www.britishcopyright.org/wp-content/uploads/BCC-Fair...

> The government should a) legislate and b) create test cases and run them through the courts so that we can have clarity.

if so, it would be nice if they approached the instances of stealing shit chronologically. but that's just my view.

reply
It's hilarious to see how all these HN intellectuals, who on any other day are rabidly against most DRM / IP issues all of a sudden become private property absolutists.

It's also sadly hypocritical to see all this rhetoric here on this thread decrying SOTAs for using a variety of content for their inputs as somehow sTeaLINg sTufF! ...

... but Chinese SOTA foundries directly using distillation as fair game.

I don't think there is any coherence to any of these arguments - other than 'we liking big companies'. That's the only common thread.

What is more reasonable:

- There's some grounds for fair use by SOTA models to ingest content, so long as they are not reproducing it ... very roughly speaking.

- SOTA makers are producing novel works, there is value add in that process, again roughly speaking.

- Distillation is a bit of a grey zone, producing random content as arbitrary input is one thing, but producing training sets is another. I think there's a coherent line in there somewhere, I'm not sure where it is.

reply