upvote
Just putting a clause in a publication won't prevent it from being used as training data. Information wants to be free.

The frontier LLM vendors do sell enterprise licenses which contractually guarantee that your prompts won't be used for training. (Maybe they'll secretly violate the agreement but in principle it's legally enforceable.) Scholars and universities who care about credit and attribution will either have to purchase those licenses or run their own private open-weight LLM instances.

reply
Even the $20 tier of ChatGPT has privacy settings that forbid using the user's data to be used for training. The question is, whether this setting is respected.
reply
I assume the data is laundered into a format that qualifies as no longer being the "user's" data, then trained on.
reply
the existence of the triplets NSA/CIA/GRU implies an imperitive no.
reply
that is a whole different thing though... AI labs are not government intelligence agencies
reply
minus two or three things, corporate data is classified as public/goverment data. we just saw something about earmarking domain last month? two being imminent domain. three being natsec.

risk of prescient theory is more important than dismissive ablation.

edit: to wit, facebook google and anything else not e2e.

it's not like the ai is homomorphic.

reply
oh, haha. another point to make: what do you think they've been building out for 25 years with fusion centers and maryland/utah?

"government intelligence agencies" ARE 'AI'.

reply
I don't like this and I wish it weren't true, but I think the period of "information wants to be free" is coming to an end, it was a relic of a bygone era. Increasingly, making your information free means you're the sucker who is doing free labor for AI companies, or worse, you're helping your competitors. Paywalls, login walls, and rate-limits are going up everywhere: there's the GitLab news on the home page right now, and sites like Twitter, Reddit etc. which used to be publicly-readable are now gated (and Xitter is using the legal system to shut down any bypasses).

I hate this but I don't think there's any going back now that LLMs exist.

reply
"Information wants to be free" never meant that people want to release their information; it meant that information is very hard to keep secret, and that everything leaks like a sieve, and especailly that once it's out, it's out forever.
reply
Exactly. While there are a few academics who work in private for years and then surprise the world with an amazing breakthrough, most of modern science and mathematics is a collaborate process. Researchers make gradual progress on hard problems, and discuss issues with colleagues and students along the way. Some of those collaborators will then pass on the information to social media or public discussion forums or free-tier LLM prompts or whatever and it gets incorporated into the next round of training runs.

Two people can keep a secret if one of them is dead.

reply
Would this legal framework cut both ways? When AI companies use AI to make and publish mathematical discoveries, would they be able to legally prevent professional mathematicians from using them?
reply
The USA doesn't have legal frameworks any more, you just buy and sell the right to do what you want. Even our supreme court is disingenuous now.
reply