You have a fundamental misunderstanding of how the watermarking works. There is no hidden metadata, but most likely, the statistical properties of the generated text are being changed.
If you met someone on Halloween with a face mask, and they’d always use the word cromulent in every sentence, you’d probably assume it’s your buddy Mark, who is about the only guy in your circle of friends who does that.
The model will produce a text where individual words at certain positions, or various n-grams encode a kind of fingerprint that will be an indicator for the text being processed by Claude.
Like, when people have, like, a specific accent or talk in a certain way, you can totally, like, figure out where they’re from, for sure.
