As AI-generated content becomes harder to distinguish from authentic media, invisible watermarks are emerging as a key way to trace how digital content was created.
Anthropic recently announced that its Claude models will begin marking text with invisible watermarks, while Google has made its visible Gemini watermark on generated images optional. The two approaches show how differently AI companies can handle content identification.
AI watermarking generally embeds information into generated media that humans cannot easily see or hear but that specialised tools can detect. For images, tiny changes to pixels can create a digital signature. Google’s SynthID, for example, spreads an invisible signature across generated images, allowing parts of the watermark to remain detectable even after cropping.
Audio can also contain signals outside the normal range of human hearing. Video may combine image and audio watermarking, although fake content can mix AI-generated and authentic elements.
For AI-generated text, watermarking works differently. Language models assign probability scores to possible words. A watermark can increase the likelihood of selected groups of words, creating a pattern that detection tools may identify. This method generally works better with longer texts.
C2PA metadata provides another way to track a file’s origin and editing history, although it is not technically a watermark.
However, detecting AI watermarks remains difficult. OpenAI has a tool for checking SynthID and C2PA, while Google offers verification features through Gemini and its own services. Google’s dedicated SynthID Detector remains invite-only for journalists and verification professionals.
Watermarks can also be removed or bypassed. SynthID is designed to withstand common modifications, while C2PA metadata can disappear when someone takes a screenshot. Text watermarks can potentially be removed by rewriting content manually or using another AI system.
Importantly, a watermark does not automatically prove that an entire piece of media is AI-generated. Authentic content can receive a watermark after AI-based editing.
The absence of a watermark also does not prove authenticity. Ultimately, AI watermark detection should be treated as one verification tool rather than definitive proof.
Do you think AI watermarks will become reliable enough to help people identify manipulated content online?


