> the legal protections for the consumer is much higher
You know you give away the right to file class action law suits against Anthropic when you accept their Terms Of Use, right? (at least the Americans ones)
I get this cynical conspiratorial energy, it fits the internet well, but I can assure you most people with even mild business sense would be intensely opposed to this idea. Well, except maybe Zuckerburg, but they don't really do enterprise anyway.
It wasn't "catastrophic" for the largest of the 3 US credit reporting agencies when their entire dataset was breached. The company is 100% IP and the only value they have was completely copied. Their largest value is to verify identities by the things Americans know (KDB) and after that "single factor of identity" was 100% compromised, the company only got bigger and more contracts.
When there are only 4 competitors in the large scale foundation model business and they all throw caution to the wind because they are racing to own the "$30 trillion TAM" they are all going to make critical security, RBAC, and segregation mistakes.
Both ChatGPT and Claude threads marked for sharing have been indexed in Google at large scale. This is incredibly easy to tell Google crawlers via robots.txt not to crawl those URLs, but nobody at either of these uber unicorns could be bothered to add that one pattern to the one file.
And all of the skepticism here is about verifiability. The foundation model companies are liable for potentially more the companies are worth if found to be violating copyrights of content used for training. They aren't going to make it easier for lawsuits against them by detailing their data ingestion into training pipeline.
But I can guarantee you that if a companies internal data got leaked or misused, then every single enterprise customer of that lab would turn around and start suing them. As an enterprise customer you would be foolish not to, if only if figure out via discovery just how badly you got screwed.
You want to see how nasty that can get. Just go and look at what Apple is doing to OpenAI at the moment. Do you really think Apple wouldn’t find a way to sue a lab into oblivion if they discovered a lab had secretly started training on their data?
1. Ensure security barriers are weak or honor based.
2. Put individual researchers under a lot of pressure.
3. If you get caught, blame the weak barriers, or the individual researcher.
Basically setup the incentive structure to incentivize researchers sticking their mittens in the private cookie jar while putting the cookie jar in a dark unmonitored/unsecured room with a sign on the door saying please don't enter.by this logic every business contract in tech is just a bunch of lies and means nothing and the only way to do anything is to have a server sitting next to you, otherwise it's "someone else's computer"
At this point, the reputation of the business matters too. Fable explicitly didn't support zero-retention usage, and it saw significantly lower adoption vs other flagship models, and their past releases. Being caught abusing enterprise contracts is really hard to dig out of.
Instead, confidentiality includes not literally copying material, not using trade secrets or inventions, and not using knowledge of business dealings for your own purposes.
So, while I fully expect that the big labs don't train on private material to the extent that they do those things in a blatant way, it would not be surprising if they pushed the boundaries. Humans push the boundaries all the time.
Up until now, machines did not have judgement, so if you set up a machine in such a way that you hadn't ensured it couldn't violate contract, you were culpable. But now that they have some kind of judgement, maybe it's enough to avoid liability to tell it to obey the contract, even if you give it incentives not to. After all, that's how it works with human employees, isn't it?
Perhaps now we have machines that understand language, someone somewhere is working on getting them to understand "a nod and a wink" as well.
(And they totally won't do it again they swear, the contract says so)
1. Child porn
2. Stolen music
3. Private github repos, before that was 'stopped'
4. Illegally pirated books
Them training on company prompts against the terms of service would be one of the least bad things that these companies have trained AI models on
Why do you think a company - willing to break the law for child porn - won't break the law when it comes to your personal data?
Have been having a think about this statement. I agree in intent, some of them are probably breaking the agreement for training data. I dont think Microsoft is doing it, Enterprise Data Protection is the plank holding up their entire Copilot line. One whiff and everyone's gone. Copilot isnt actually good at anything except giving some illusion of protection, and preventing users from following a desire path to other LLMs without enterprise data protection. Its the core value proposition. But we only need to wait and see what the next 20 data breaches tell us to find out for sure.
In detail however, I dont know if they could just claim it as fair use, after exclaiming that they specifically wont do that. I dont think "Fair Use" would be the issue so much as contract law. I know a EULA wouldnt hold up the other way (By reading this you agree not to steal my data and train an LLM with it) for data thats made freely available on the internet. But if you are purchasing the "No Training" contract they would be in breach if they trained with it. Possibly fair use would let them keep the data after paying whatever they owe in terms of contract breach.
And yet, software companies are donating their means of production to these crooks left and right, hastening the day where they really will become obsolete. I'm sure others do equally unwise things.
If you're big enough, crime pays exceedingly well.