upvote
My parents put in countless hours and tens of thousands of dollars into raising me to the point where I could write an answer on StackOverflow

And OpenAI scraped and distilled that answer and gave me nothing

reply
And now people such as myself have access to open weight models with that information. I wasn't lucky enough to have parents put me through school, and LLMs have absolutely helped me further educate myself and play "catch up" on opportunities others have been given. So, the net effect has been (and is continuing to be) a democratization of information.
reply
You could have gained that stuff prior to LLMs. The leg up you're describing is free information on the internet, not AI. AI just makes it a little easier to find, while also crushing the original sources in the process. (Even if it had a broken culture, is stack overflow even going to exist in a year? Where are they going to train on going forward?)
reply
Democratization of information, but Sam Altman gets a $100B net worth and I'm still broke :)

I would prefer some sort of democratiziation of the money made from the democratization of information as well

reply
Not to be an arse, but didn't you have access to Stack Overflow with all questions/answers prior to LLMs?
reply
Isn’t that the same argument they are making for replacing human labour?

Circumventing costs.

reply
There are many frames that one can place upon this issue. They do not contradict the other. There are moral framings (stole the internet so go eff yourselves, is one), but so is national security, and so is the doomer recursive self improvement risk, and then there is the framing purely on what this implies for future AI training.

I mainly focus on the last.

It will be hard for a frontier lab to justify spending the compute and data curation needed to advance AI further if that expenditure can be assimilated into your competitor's products within months/weeks. So reality will present labs with three choices:

A. Cease spending massive amounts of money and compute improving those models.

B. make those improved models more difficult to distill from, either through some regulatory regime, or some technical solution, which seems unlikely to me.

C. making the best models available only to select partners and government.

In all these potential outcomes, China, which lacks compute that U.S. labs enjoy, will likely stop seeing massive improvements in their AI models. Improvements to be sure, but right now they are enjoying gains from distillation AND their own model innovations, and these potential outcomes would largely stop one of those sources.

reply
Oh, so mass theft is okay as long as American companies are doing it
reply
copyright infringement is not theft, even if right holders often claim it is.

part of the definition of theft is that the original owner is deprived of it, which does not apply to copyright infringement.

You can only argue with damages from the perspective of potential profits, still not theft though.

https://en.wikipedia.org/wiki/Theft

reply
So having tons of AIs quoting various literary works and reproducing knock-offs of them has a positive effect on those books' sales?

I think you're wrong: there is absolutely damage to the authors and publishers from what the AI companies have done.

reply
If you steal an unpopular product from a store, the damage is also only to "potential profits", so how does that differ? It's entirely possible no one would have purchased the product and it would have eventually been discarded/destroyed.

Or with services, if a barber cuts your hair and then you run away without paying them, do you not consider that theft, even though there's no change in ownership occurring?

reply
Moreover, reading a copyrighted book and learning from it is not theft.
reply
Generating a set of weights is not learning.
reply
would you say that airplanes don't fly because they don't flap their wings? it's possible to achieve the same things with different approaches.
reply
That is a strong statement. I guess you are telling Machine Learning to go fuck itself.
reply
No, Machine Learning is an unfortunate name for a well documented process for creating black-box classifiers. The process is good, the name is not.
reply
And what, to your mind, would classify something as learning? I assume that your position is not the hard "only humans/living creatures can learn"
reply
Great! Neither is distillation then.
reply
Never said it was. Still, understanding to what extent the ability of Chinese labs to keep up to western models with much less compute needs to be understood.
reply
Machines are not humans.
reply
Reread my comment and look for a value judgement on my part. The final sentence is probably a good clue as to my opinion.
reply
Chatgpt routinely cites and uses papers I don't have access to because they're behind a paywall. I don't think OpenAI is paying for all that copyright. That's in my opinion way more serious.
reply
Yes the fact that the scientific literature - created largely on the back of the tax payer - isn't open to all free of charge by force of law is a travesty. A cartel should not get to charge for access to the bulk of human knowledge. That is indeed a far more important issue than whether or not Moonshot violated the Anthropic ToS, possibly committing mass fraud in the course of doing so.

I mean honestly if they did that why should I care? I'm happy to see copyright violated in a manner that leads to the creation of new technology. IP law exists strictly for the benefit of society and by all appearances AI is an incredibly powerful tool.

Also while I'm at it libgen is a gift to humanity. Information wants to be free. Spreading and preserving knowledge is generally one of the most wholesome activities anyone can undertake as far as I'm concerned.

reply
Your statement is orthogonal to my comment. Why reiterate the schadenfreude / fairness comment already stated several dozen times in this thread?
reply
Who do you think is paying $100K+ for "Enterprise" access to Anna's Archive?
reply