General_Effort

@General_Effort@lemmy.world · Joined ⁨Dec⁩ ⁨2023⁩

Replying to @⁨Prox@lemmy.world⁩

You know how these AIs output tokens? I’m going to explain the concept with words, because that makes it easier to understand. But it means that the explanation is [ETA:]NOT quite right.

An AI has a vocabulary. The words in that vocabulary are assigned to 2 groups.

Sometimes when the AI is outputting something, it could use different words equally well. It’s not quite the same as having synonyms, cause this isn’t really about words. But let’s say you have synonyms in those different groups. At those points in the text, you can pick from one group or the other to embed a hidden pattern in the text.

Limitations are obvious. To embed the watermark, you need enough opportunities to pick “synonyms”. It won’t work for very short texts, or if the word choices are very constrained.

I’m curious if the negative effects are really as minor as they say.

Replying to @⁨Goodlucksil@lemmy.dbzer0.com⁩

Copyright. They mustn’t make a copy. There is legal precedent that confirms it’s okay to transform the copy you bought into digital format. The Internet Archive relies a lot on that. I think they actually litigated it in the first place. So that’s why the copyright heads are going so absolutely apeshit. If no one’s charging you rent for using some data, then it’s “unethical”.