theeyeinside
Senior Member
- Aug 19, 2014
- 962
- 867
https://www.cyberseo.net/blog/google-geminis-synthid-a-hidden-trap-for-autoblogging/
In text, this is not done through metadata or hidden symbols, but rather, at the level of token probabilities. When a language model selects the next word, it uses a probability distribution. For example, the probabilities for “cat,” “dog,” and “bird” are 0.31, 0.28, and 0.14, respectively. SynthID shifts these probabilities slightly according to a pseudorandom rule.
The result is text that remains readable and natural to humans but that, when analyzed for word sequences (e.g., n-grams of five tokens), exhibits a distinctive statistical pattern. This pattern is invisible to the reader, but Google’s detector can identify it: “This text was generated and watermarked by SynthID.”
The principle is similar for images. During generation, tiny distortions are embedded at the pixel and frequency levels. These distortions are imperceptible to the human eye, yet they survive compression, resizing, cropping, and filtering. With its detector, Google can easily identify images created by its model.
If you want to dive deeper into the technical details, check out the Introducing SynthID Text.
---------
Eventually all the models will do this and AI content will be easily uncovered
Writers will rejoice and not lose their jobs after all ...
How SynthID works
In text, this is not done through metadata or hidden symbols, but rather, at the level of token probabilities. When a language model selects the next word, it uses a probability distribution. For example, the probabilities for “cat,” “dog,” and “bird” are 0.31, 0.28, and 0.14, respectively. SynthID shifts these probabilities slightly according to a pseudorandom rule.
The result is text that remains readable and natural to humans but that, when analyzed for word sequences (e.g., n-grams of five tokens), exhibits a distinctive statistical pattern. This pattern is invisible to the reader, but Google’s detector can identify it: “This text was generated and watermarked by SynthID.”
The principle is similar for images. During generation, tiny distortions are embedded at the pixel and frequency levels. These distortions are imperceptible to the human eye, yet they survive compression, resizing, cropping, and filtering. With its detector, Google can easily identify images created by its model.
If you want to dive deeper into the technical details, check out the Introducing SynthID Text.
---------
Eventually all the models will do this and AI content will be easily uncovered
Writers will rejoice and not lose their jobs after all ...