upvote
They still need to choose when to do that though. When I prompt the program to e.g. alter a bash script in a specific way or to recite a longer known text it can't go round and randomly exchange tokens. It has to somehow define what is a simple repeated text from a different origin and what is a novel generation.
reply
I am wondering how that applies to newly generated code.

Odd variable naming? Stylistic choices that are watermarked?

Or as someone else noted further down in the comments, it could be more subtle:

Between the first and second most likely choice, in certain positions it will consistently choose in a certain way.

reply
> Odd variable naming? Stylistic choices that are watermarked?

Whatever it is, I'm sure it's load-bearing.

reply
You're absolutely right. But it is not just load-bearing, it is the load-bearing seams.
reply
I would guess they're not worrying about watermarking a tweak to a human-written program. That's both a tiny fraction of Claude use and of very little concern to the kinds of people who want to check watermarks.
reply
Most probable usually means for a specific prompt. How can this operate without the the original prompt?
reply
Just double checking my understanding: If this is true then only Anthropic will be able to detect if text was generated by one of its models, correct?
reply
Likely yes.
reply
But what prevents someone from using Anthropic own detection system to train a watermark-scrubber?

Seems like this would only catch the most unsophisticated cases.

reply