upvote
There exist methods to detect what kind of watermarking tool that someone is using, and most of the big tools have specific signatures that you can look for.

For Anthropic, it’s highly likely that the watermark is a SynthID type mark similar to the one that Sean is talking about (I actually ran the analysis here https://johnjwang.com/post/2026/08/12/how-claude-watermarkin...). When we get confirmation of whether all models actually are watermarked, I think we’ll be even more confident.

Of course it’s always possible that Anthropic has come up with a proprietary scheme, but I think it’s definitely harder to implement.

I think the game will be a cat and mouse game similar to LinkedIn and other websites trying to block scrapers: each iteration makes it harder for someone to figure out the watermarking scheme, but likely not impossible

reply
I don't know if it's possible to achieve a watermark that is undetectable, has a low false positive rate, and survives a wholesale rephrasing. Can you make a statistical measure that reliably survives 95+% of the words being different and the sentences reordered? Of course the more of the content you replace, the lower the quality, but in many cases you probably care more about the meaning of the text than the exact choice of words.
reply
>...and survives a wholesale rephrasing.

There's also literal language translation. Generate in language A, translate to language B (Either "manually" by being proficient in it, or with non-LLM translation).

A very, very significant portion of the world knows more than 1 language.

reply
If you can submit the text to determine whether it's watermarked, you just have to progressively alter the content more and more until it passes.
reply
It could be, but common, we all know that Anthropic's watermark is the using "load bearing", "genuine", and "seam" 1000x more in the same paragraph than any human in history.
reply
[dead]
reply