upvote
Well, for a start, they’re probably on dodgy ground there with other European regulations. You’re not really supposed to hold onto hold onto potentially very sensitive data that you have no legitimate interest (a term of art; it doesn’t just mean “I want to keep this”) in indefinitely.

As far as I know, most of these require user consent for retention?

reply
> AI companies already store all prompts and responses for future training.

They store some prompts and responses, not all, that's what you're missing.

reply
> AI companies already store all prompts and responses for future training.

They claim to only do this when you agree to this in your personal settings. Though Google does say they will train on it, unless you disable history and only use ephemeral chats. Anthropic has a setting for it and claims not to train by default.

Also it would be vulnerable to attacks and privacy problems. You could search for substrings about some suspected information, like "John Smith's medical records show advanced cancer" etc. Of course you'd have to guess the phrasing but still.

reply
Wouldn't this become pretty difficult to do at scale over time? Is there a way to compare the similarity in a database of responses without doing a search over every entry and comparing them? Because that would probably become pretty slow if literally every LLM output is saved and has to be scanned.
reply
1. AI company buys and trains on an author’s book when it gets published, it’s now part of the training data.

2. Attacker asks the LLM for the opening sentences of the book, it goes into the generated responses database.

3. Later, a malicious user shows that the first few sentences of the authors book are identical to a previously generated response.

reply
Ignoring the other technical hurdles of the idea...this problem you outlaid is solved with a timestamp in the database, right? You can easily prove if the prompt was before/after publishing date?
reply
Local models?
reply
Local models enable:

- watermark-free generation

- the stripping of watermarking from the output of SAAS models

Any discussion of watermarking is dead in the water in a world where we are permitted to have these things. I fear for the future.

reply
Local models also enable:

- not being locked into a provider

- not being forced to have your prompts saved by a possible competitor

- an alternative to the duolopy we quickly see forming

- offline access

I fear for the future without local models, much more than the future with them, and would rather everyone had access to a local model than be certain we catch everyone copy+pasting LLM responses. Watermarking would be cool, but it's not worth losing local for.

reply
To be clear, we agree. The problem is that unless local AIs become "normie friendly" real damn quick, we're gonna lose em, because they're damned inconvenient to power. That is what I fear.
reply
My apologies, I misunderstood your original comment to mean "We shouldn't have local models if it breaks watermarking" but yes, I think we're on the same page now
reply
Chill my dude. This is just a sane default which will catch normies copy pasting stuff from claude and chatgpt. It's good enough.
reply
Do you not want "normies" to have local AI? Do you believe that your ability to use esoteric software makes you special, makes you immune to law? The days where you can find refuge in running "weird hacker crap" instead of 100% goobermint approved software are numbered because news flash: AI means there are no normies anymore. Anyone can do anything.

No, I will not chill. The war on general purpose computing is gonna get real hot real soon.

reply