Wouldn't this make it a worse measurement?
Teaching you how to identify watermarked text while the experiment is running would ruin the data.
Best practice is to allow a number (scaled based on complexity of task) of training rounds (with short feedback loops) prior to letting people loose on the regular samples.